Claude Ran a 31M-Token Math Sprint and Broke a 6-Year Record on the Riemann Hypothesis
Honestly! AI's breakthroughs in the mathematical world are becoming more and more casual.
The trigger was quite casual. Eight days ago, Anthropic employee Jarred Sumner, while running, posed an "unreasonable challenge" to an unreleased research version of Claude: seriously tackle the Riemann Hypothesis.
Claude's reasoning process went through about 650 failed ideas, 60 sub-agent relays, and a day and a half of continuous derivation. Although it didn't prove the Riemann Hypothesis, the model incidentally raised a lower-bound record directly related to the Riemann Hypothesis from 41.6% to 67.2%.
On August 11, Anthropic publicly released the paper, process records, and formal proof on its official website.
Surprisingly, Anthropic employee Jarred Sumner is not a mathematician; he dropped out of school at 16. His interactions with Claude were mostly encouraging messages like "keep pushing" and "believe in yourself."
This made me curious: just how difficult is the Riemann Hypothesis? And what does raising this lower bound actually mean?
◈ Stuck for 167 Years, Claude Raises the Lower Bound
The Riemann Hypothesis is one of the most famous unsolved problems in mathematics.
The Clay Mathematics Institute listed it among the seven Millennium Prize Problems, with a $1 million reward. To date, only one problem has been solved.
The Riemann Hypothesis was proposed in 1859, making it 167 years old today. Countless mathematicians have tried and failed to solve it in that time.
Before understanding the Riemann Hypothesis, let's first grasp these terms.
Prime numbers seem to appear at irregular intervals: after 2, 3, 5, 7, we jump to 11, then to 13.
The Riemann zeta function encodes information about the distribution of prime numbers into a complex function. The special positions where this function equals 0 are called "zeros."
The ones that truly concern the Riemann Hypothesis are called non-trivial zeros.
The Riemann Hypothesis states: none of these non-trivial zeros can deviate; they must all lie exactly on the critical line where the real part is 1/2.
It's as if mathematicians stretched an extremely thin wire across the boundless night sky and then asserted: all the key stars are pinned to this line.
Mathematicians still cannot prove "all."
So they took a step back: at least what proportion of zeros can be proven to lie on the line?
In 1914, Hardy proved there are infinitely many zeros on the critical line. In 1942, Selberg proved that a positive proportion of zeros lie on the line. In 1974, Levinson pushed the explicit lower bound to one-third. In 1989, Brian Conrey pushed it past two-fifths. By 2020, Kyle Pratt, Nicolas Robles, Alexandru Zaharescu, and Dirk Zeindler pushed the record to slightly above 5/12, roughly 41.67%.
Then, this number stalled for about 6 years.
Until this time, Claude produced a result of 67.25%. Compared to the previous record, it jumped by about 25.58 percentage points in one go.
Of course, 67.2% is not the "completion rate" of the Riemann Hypothesis. The remaining 32.8% is simply the portion this argument has not yet classified, not that they are off the critical line. Anthropic also explicitly stated they do not expect this set of techniques to lead directly to 100%.
◈ Claude's Wild Reasoning
When Jarred first asked Claude to try, the model generated and tested about 650 ideas in one go.
Not a single one worked.
Normally, this would be the end. Claude itself knew how difficult the Riemann Hypothesis is and once doubted it could make any meaningful progress.
Jarred didn't give up; he just told it to try again.
In the second round, Claude no longer worked like a person sitting at a desk struggling alone. It coordinated about 60 sub-agents within Claude Code, splitting the task into finding ideas, doing calculations, checking counterexamples, reviewing proofs, and writing the paper.
Of the 60 sub-agents, only 2 developed key mathematical ideas; 13 provided auxiliary ideas; 30 tried new directions but didn't succeed; 13 were responsible for verification; and the last 2 helped organize the initial paper.
Half the sub-agents failed.
Over a day and a half, this Claude team ran about 2,400 shell commands, wrote hundreds of Python scripts, and performed thousands of numerical checks on known zeta zeros. After finding the result, they continued searching for counterexamples, had different sub-agents re-prove from scratch, and even downloaded 54 arXiv papers to check if related results already existed.
The two sessions combined consumed 31 million output tokens.
Finally, Claude proactively wrote the results into a paper.
Jarred barely intervened in this round; his involvement was reduced to sending encouraging messages: "Continue," "Believe in yourself."
◈ As Soon as the Paper Was Submitted, Four Mathematicians Took Over
The one thing open problems never lack is "I proved it."
Anthropic's internal mathematicians Levent Alpöge and Ralph Furman studied and verified Claude's paper, and also wrote a simplified proof explanation for human experts.
Two external number theory experts, Brian Conrey and Dan Goldston, also quickly reviewed the results after receiving the paper on short notice.
Conrey had previously pushed the related lower bound past two-fifths in 1989; Goldston participated in the prior research that Claude relied on this time.
Claude also collaborated with Anthropic employee Eric Easley to write the results from the paper into a Lean formal proof. The public repository states that this code left no sorry placeholders for unproven statements and passed the standard verification tool comparator.
The comparator checks whether the submitted proof truly corresponds to the pre-given theorem statement, rather than quietly swapping in an easier problem.
All of this gives this achievement an evidence chain far stronger than an ordinary AI demo.
This breakthrough wasn't conjured out of thin air by Claude.
Jarred himself wrote clearly in his post: "I made no mathematical contribution to this paper," with the greater credit going to Claude.
Claude's reasoning process heavily relied on recent research on zero spacing by Baluyot, Goldston, Suriajaya, and Turnage-Butterbaugh, as well as Bombieri's 2000 paper. These human results provided the necessary statistical information and analytical tools.
AI's progress in mathematics is truly getting faster and faster.
In a day and a half, Claude raised a lower bound that had been stuck for about 6 years from 41.6% to 67.2%.
But at this pace, AI won't stay forever on the periphery of difficult mathematical problems. Next time, we might see AI not just incidentally refresh a lower bound, but push a truly difficult scientific problem past the finish line.
Claude Opus 5 is here! Five ways to use it in China (2026 latest)
Top 1 of 2 from juejin.cn, machine-translated. The original thread is authoritative.
Is this because they feel the company gives them too many tokens and they want to hit a KPI?
Their quota is probably unlimited, after all, 80% profit [crying]