Claude Ran a 31M-Token Math Sprint and Broke a 6-Year Record on the Riemann Hypothesis
A frontier model, given a hard open problem and left to run, produced a mathematically meaningful result that four domain experts verified and that survives formal proof checking. The jump from 41.6% to 67.2% doesn't solve the Riemann Hypothesis, but it shows AI can now advance a stalled mathematical record when the right prior literature exists to build on.
Anthropic let a research Claude loose on the Riemann Hypothesis with minimal human steering. After 650 dead-end ideas in the first attempt, a second run coordinated roughly 60 sub-agents across a day and a half, burning 31 million output tokens. The model didn't prove the hypothesis, but it raised a related lower bound by over 25 percentage points — the first movement on that metric in six years. The result was verified by four mathematicians, including Brian Conrey, who set the prior two-fifths record in 1989, and the proof was formalized in Lean with no `sorry` placeholders. The human in the loop, Jarred Sumner, is a 16-year-old dropout with no math background; his contributions were limited to messages like "keep going" and "believe in yourself."
The result's credibility rests less on the AI's output and more on the verification pipeline: four human experts plus a Lean formal proof that passes comparator checks. That's a higher bar than most AI math demos clear.
Half the sub-agents failed, and only two produced key ideas — a reminder that scaling agent count doesn't linearly produce insight. The architecture succeeded by generating enough shots that a few stuck.
The model's dependence on recent human papers (Baluyot, Goldston, Suriajaya, Turnage-Butterbaugh, Bombieri) suggests the breakthrough was combinatorial assembly of existing results rather than novel theory generation. That's still useful, but it bounds the kind of problem this approach can currently crack.
Anthropic's public release of the paper, process log, and formal proof sets a transparency standard that makes independent verification possible — a contrast to the typical AI demo that ships a cherry-picked output with no audit trail.
Is this because they feel the company gives them too many tokens and they want to hit a KPI?
Their quota is probably unlimited, after all, 80% profit [crying]