⚡ Uncle Cat AI Radar
ResearchAgents

Claude lifts Riemann zeta lower bound to 67.2%

Anthropic says an unreleased research Claude raised the proven share of zeta zeros on the critical line from 41.6% to 67.2%, checked by outside experts.

Anthropic on Monday published an account of asking an unreleased research version of Claude to attempt the Riemann hypothesis. The model did not prove it. It did produce a new result on a well-known adjacent problem.

What the result is

Mathematicians have long sought to bound the proportion of non-trivial zeros of the Riemann zeta function that lie on the critical line — the property the hypothesis asserts holds for all of them. The prior published state of the art stood at about 41.6%. Anthropic says Claude raised that lower bound to 67.2%, by recognising that results from Baluyot, Goldston, Suriajaya and Turnage-Butterbaugh could be combined with earlier work by Bombieri to clear the previous ceiling.

How it was produced and checked

The run was not a single prompt. Anthropic describes two sessions in Claude Code consuming roughly 31 million output tokens: the model generated about 650 candidate ideas, then orchestrated some 60 subagents that executed around 2,400 shell commands and wrote hundreds of Python scripts to numerically verify claims, hunt counterexamples, re-prove steps independently and critique the draft.

Two mathematicians employed by Anthropic, Levent Alpöge and Ralph Furman, validated the output, and analytic number theory specialists Brian Conrey and Dan Goldston reviewed the paper on short notice. Anthropic states plainly that it does not expect these techniques to lead to a proof of the hypothesis itself.

Why it matters

This is a concrete, expert-inspected advance on a problem professional number theorists have pushed at for decades, not a benchmark score. But the caveats are structural: the work has not been through conventional peer review, and because the model is unreleased and unnamed, no outside party can reproduce the run. The industry is now generating research claims faster than the verification infrastructure around them exists — which makes agent-scale orchestration, and the 31-million-token compute bill behind it, as much of the story as the mathematics.

Sources