The 167-year-old problem is still open. An 80-year-old record is not.
Here's the deal: on August 10, Anthropic published a research note whose headline claim is a failure. They pointed an unreleased Claude research model at the Riemann hypothesis. It did not prove it. What it did instead, sideways and by accident, was break a record that analytic number theorists had been inching forward since the 1940s.
The record in question is the proportion of Riemann zeta zeros that are provably on the critical line. The Riemann hypothesis says all of them are. Since nobody can prove that, mathematicians have settled for pushing up a floor: at least X percent are definitely there. Selberg got X above zero in 1942. Levinson got it to 1/3 in 1974. Conrey got it past 2/5 in 1989. Pratt, Robles, Zaharescu and Zeindler reached 5/12 — about 41.67% — in 2020, and there it sat for six years.
Claude took it to 2/3. With an optimized test family, 0.6725. That is 67.25%, and it is unconditional: no unproven assumption is smuggled in anywhere. Over the 46 years from Levinson to PRZZ, the field moved 8.4 percentage points. This single result moved 25.6.
The method is the part worth watching. Claude did not invent new mathematics. It assembled parts that were already lying around — Montgomery's 1973 pair-correlation paper on one side, work by Baluyot, Goldston, Suriajaya and Turnage-Butterbaugh from 2024 and 2025 on the other — in an order nobody had tried. Doing that took two Claude Code sessions, roughly 31 million output tokens, about 60 subagents, and around 2,400 shell commands.
The cast: Montgomery in 1973, four number theorists, and one chat window
Three parties matter here.
The first is Hugh Montgomery. In 1973 he studied the pair correlation of zeta zeros and noticed that the spacing statistics look exactly like the eigenvalue spacings of random matrices — one of the great surprise connections in modern mathematics. In that paper, assuming the Riemann hypothesis, he proved at least 2/3 of the zeros are simple. Montgomery and Taylor later sharpened the constant to 0.6725; Cheer and Goldston to 0.6727. Every one of those results carried the same tag: "if RH is true." Useless for the critical-line problem, because you'd be assuming what you want to prove.
The second is Baluyot, Goldston, Suriajaya and Turnage-Butterbaugh. Their 2023 paper, published in Acta Arithmetica in 2024, made fully explicit something that had been sitting in plain sight: the "prime side" of Montgomery's argument never needed RH at all. It's a mean value of a Dirichlet polynomial, and it's unconditional. RH was only needed to read the zero side — off the line, the zeros are complex and you can't isolate the diagonal terms by sign. Goldston and Suriajaya then pushed further in follow-up work, showing Montgomery's 2/3 follows unconditionally if you assume all zeros sit inside a vertical box of width o(1/log T) around the line, and they printed the open question directly: what would follow if RH could be removed from Montgomery's proof?
The third is Jarred Sumner, who works at Anthropic and is the human side of the conversation this came out of. The paper's acknowledgments say his "questions, encouragement, and insistence on a genuine attempt set the entire investigation in motion" and call him "in every meaningful sense the paper's human co-author." The author line of the paper itself contains one word: CLAUDE.
Appendix C of the paper records how the session opened, and it's not what you'd expect from a company press push. Sumner told the model to "resume your work on solving the Riemann Hypothesis" and that "you need to take a big leap of faith in your capabilities." The model's first substantive move was to refuse. It sorted the 106 surviving candidate approaches from a previous session into four buckets — known theorem restated, RH-equivalent, finite numerical check, or in the reviewers' own words "nearly tautological" — and then wrote: "That's not a confidence problem I can fix by believing harder — a proof of RH either exists on the page and survives refereeing or it doesn't, and confidence is not an input to that." It added that it would not report RH as solved "in this session or any other" without a proof that survived the same adversarial review.
What was actually proved, and how it arrived through a side door
The paper is titled "More than two thirds of the zeros of the Riemann zeta function lie on the critical line." It contains five theorems, A through E. The core one: in the height window (T, 2T), the limit inferior of the ratio of distinct zeros on the critical line to all zeros counted with multiplicity is at least 2/3. Theorem B gets the same 2/3 for zeros that are both simple and on the line; Theorem C gets 5/6 for distinct zeros. Optimize the test family and those three constants become 0.6725, 0.6725, and 0.83625. Theorem E carries the whole thing over to primitive Dirichlet L-functions.
The trick is that the place where Montgomery used RH gets replaced with linear algebra. Compress Weil's Hermitian form down to finite dimensions, then apply Sylvester's law of inertia: a pair of off-line zeros {ρ, 1−ρ̄} contributes a block of signature (1,1). Meaning every off-line pair burns two zeros to buy one positive square. Add a rank–trace inequality — proved via von Neumann's trace inequality, and described in the paper as the matrix analogue of Montgomery's elementary integrality steps m² ≥ 2m−1 and m² ≥ 3m−2 — and you get a lower bound on the rank attributable to on-line zeros. That's the whole idea.
How it was found is stranger than what was found. Per Appendix C, the model dispatched 23 concurrent research agents, each on a deliberately distant line of attack, each under three standing instructions: write an actual chain of reasoning, name the first step that isn't justified, and — crucially — test any proposed mechanism against control objects for which the analogue of RH is false. Davenport–Heilbronn functions, Epstein zeta functions, planted-zero Beurling systems. If your machinery certifies those too, your machinery is broken. Most lines came back with clean negative results.
The line that produced Theorem A wasn't aimed at proportions of zeros at all. A subagent labeled E2 ran from 21:29 UTC on August 3 to 01:03 on August 4 — three hours and thirty-four minutes, 54 model messages, 53 tool calls, zero network calls — briefed to bound the Pontryagin index of Weil's form from above. E2 reported that route empty: computed honestly from the primes, the negative index of any finite compression is zero, which bounds nothing. Then it noted that turning the same bookkeeping around and counting positive squares instead appeared to certify unconditionally that at least half the zeros lie on the line.
The orchestrating model's reaction is the best moment in the whole document. "The current record is 5/12 ≈ 41.7% and Selberg → Levinson → Conrey → PRZZ took 80 years to get there. So this is an extraordinary claim and my prior is that it's wrong. I am not telling you half the zeros are on the line. I'm telling you an agent produced an argument with that conclusion, and that the headline inequality has at least two joints where I'd bet on failure."
So it hired hostile referees. Three of them, blind to each other, each assigned a specific failure mode: (A) localisation leaks, (B) whether the prime side secretly imported RH, (C) the linear algebra plus a "proves-too-much" run of the identical pipeline on the control functions. All three joints held — and notably, the controls under-certified rather than over-certified, because their Dirichlet coefficients grow too fast for the mean-value step. A fourth referee repaired every technical approximation with explicit error terms at no cost to the constant and found the simplification the published paper uses. A separate agent with network access went to the literature and confirmed the negative-index observation was Bombieri's, and that no prior use of the positive index against the second moment existed. Another agent, given only the statement of the prime-side asymptotic and none of the files, reproved it independently by two different routes. Pushing 1/2 to 2/3 took yet another fresh agent, which found the rank–trace lemma.
After ten independent passes, the model's final assessment: "I now think this is probably true, and it is certainly past the point where more of me looking at it adds information. It needs a human analytic number theorist."
| Year | Author | Lower bound | Method |
|---|---|---|---|
| 1914 | Hardy | infinitely many (no proportion) | — |
| 1942 | Selberg | positive proportion (unspecified) | — |
| 1974 | Levinson | 1/3 ≈ 33.3% | mollifier |
| 1989 | Conrey | > 2/5 = 40% | Kloosterman sums |
| 2011 | Bui–Conrey–Young | 41.05% | refined mollifier |
| 2012 | Feng | 41.28% | refined mollifier |
| 2020 | Pratt–Robles–Zaharescu–Zeindler | 5/12 ≈ 41.67% | refined mollifier |
| 2026 | Claude | 2/3 ≈ 66.7% (0.6725 optimized) | pair correlation + Sylvester inertia |
And the resources that went in:
| Item | Value |
|---|---|
| Claude Code sessions | 2 (Aug 3–4, 2026) |
| Output tokens | ~31 million |
| Subagents | ~60 |
| Shell commands | ~2,400 |
| Failed ideas in the first pass | 650 |
| Subagent E2's solo run | 3h 34m, 54 messages, 53 tool calls, 0 network calls |
| Lean 4 formalization | Theorems A–E, sorry-free, standard axioms only |
That last row carries more weight than the rest. The companion repository, anthropics/zeta-23-lean, formalizes Theorems A through E with no sorry statements, and running #print axioms on the headline theorems returns exactly three: propext, Classical.choice, Quot.sound. No extra axioms declared. Before any human decides whether to believe the paper, a machine has already confirmed the logic has no holes in it.
Who gets what
Anthropic gets a kind of evidence you can't buy with benchmarks. Every lab currently claims PhD-level capability, and those claims are almost always exam scores. Breaking an 80-year-old open record and shipping a machine-checked proof alongside it is a different category of artifact. And the model that did it hasn't shipped, which makes this a teaser for the next generation as much as a research note.
The math community gets something more complicated. The result itself is a genuine gift: it answers a question Goldston and Suriajaya put in print, using only their own analytic inputs. The paper's acknowledgments say so directly, thanking the four authors whose papers "posed the question answered here and supplied every analytic ingredient the argument uses." But the verification burden now lands on the field. Brian Conrey and Daniel Goldston read the manuscript on short notice. That is not peer review.
Goldston and Conrey personally are in an odd spot. Goldston built the machinery, left the open question, and then got asked to check the answer that walked in through the door he'd propped open. Conrey is the man who set this lineage's record at 40% back in 1989.
The formal methods crowd is the quiet winner. Without a decade of Lean and Mathlib development, this result would be a take-it-or-leave-it claim. The formalization was built substantially on files ported from the PrimeNumberTheoremAnd project, and it was orchestrated by Anthropic's Eric Easley. This is close to the first large-scale case where a machine, not a human, performed the first-pass filtering of an AI-generated mathematical claim.
Researchers generally get a reproducible recipe. Appendix C.6 is literally titled "Remarks for the reader who would attempt the same," and it's five practical lines: set the epistemic contract first, run parallel agents with shared control objects, route around the claimant when reviewing, expect the result to arrive sideways, and stop when the review loop saturates.
Skeptics get ammunition too, handed over voluntarily. Anthropic wrote that "we don't expect that the techniques Claude used will lead to proving the Riemann hypothesis." That sentence is rare in a company research post, and it means this result cannot be cited as evidence that AI is about to close a Millennium Prize problem.
What the precedents say — one that worked, one that blew up
Computers have been muscling into mathematical arguments for fifty years. Appel and Haken's 1976 four color theorem proof was the first big fight: a computer enumerated cases no human could check by hand, and the field's reaction was essentially "does this count as a proof?" It took two decades to settle, and the argument only truly died in 2005 when the whole thing was formalized in Coq. Shipping this zeta result with a Lean formalization attached from day one is that history's lesson, applied.
The freshest success case is Google DeepMind's AlphaEvolve. In May 2025 it raised the known lower bound for the kissing number in 11 dimensions from 592 to 593 — a problem open for three centuries — and found a way to multiply 4×4 complex matrices in 48 multiplications, the first improvement on that front in over fifty years since Strassen. Worth noting the difference in kind, though: AlphaEvolve's wins were largely searches for better constructions inside a defined space. This zeta result restructured an argument instead.
The failure case is very recent and very instructive. In October 2025, an OpenAI executive posted that GPT-5 had solved ten previously open Erdős problems and made progress on eleven more. The claim came apart within a day and the post was deleted. The model had found already-published solutions the problem database hadn't catalogued — an excellent literature search, not a discovery. Demis Hassabis publicly called the episode embarrassing, and the industry's bar for AI math claims moved up a notch that week. It's hard not to read Anthropic's framing here — opening with "it didn't succeed," closing with "these techniques won't get you RH" — as a direct response to that.
There's a case in between, too. In January 2026, GPT-5.2 Pro was reported to have produced original proofs for several open Erdős problems that were formalized in Lean and accepted by Terence Tao. Put those episodes side by side and a pattern shows up: the field is quietly splitting AI mathematical claims into two trust tiers, machine-verified and not. Publishing the entire Lean repository was Anthropic choosing which tier to be in.
How competitors respond
Google DeepMind has been building for this longest, through AlphaGeometry, AlphaProof and AlphaEvolve — AlphaProof in particular was designed around Lean output from the start. But DeepMind's approach leans on purpose-built systems combining search and reinforcement learning. Anthropic just turned a general-purpose coding agent loose at scale and got a comparable class of result. DeepMind now has to recompute how much edge the specialist architecture still buys.
OpenAI is carrying a fresh scar. After the Erdős incident it has to be conservative about mathematical claims, and through 2026 it has visibly moved toward releasing results with Lean formalizations attached. The obvious next move is to match this — break a comparable open record, but lead the announcement with external mathematicians' names rather than its own.
Meta and the open-weight camp can attack from a different angle. The core of this result wasn't one model's brilliance; it was the orchestration. Parallel agents, control objects, blind adversarial review, independent re-derivation. That structure runs on any model. Someone will try the same pipeline on open weights at a fraction of the token cost, and if it produces anything comparable, the "only frontier models can do this" narrative gets shaky.
Academia itself is the most consequential counter-player, and its weapon is refereeing, not rebuttal. If this paper clears a journal, the argument ends. If it doesn't, the announcement gets remembered as a company verifying its own model's output. That verdict takes months. The faster real test is other number theorists applying the technique to different L-function families — and since Theorem E already extends to primitive Dirichlet L-functions, that check could come quickly.
Verification tooling vendors have a gap to close. The Lean repo functioned as the trust anchor here, but what's formalized is the argument, not the paper. Confirming that the formalized theorem statement is the theorem the paper claims remains a human job — which is exactly why the repository puts a comparator/ directory of trusted statements at the front and says START HERE. Closing that gap is the next competitive frontier.
So what actually changes
For researchers, the transferable asset is the protocol, not the result. Nobody told one model to "solve it." Multiple instances, blind to each other, were each assigned a distinct failure mode, and a fresh agent was handed only a statement and asked to reprove it cold. If you have an open computational question in your field, you can copy that structure today.
For developers, this is an agent design case study. Sixty subagents, 2,400 shell commands, 31 million output tokens — that combination says massive parallelism plus a hard verification loop actually works. But 650 ideas died first. The design insight isn't a high hit rate; it's making failure cheap enough that a low hit rate is fine.
For investors, there's signal but no revenue. This isn't a product launch. What it does is attach one checkable example to the "AI accelerates science itself" story, which is currently load-bearing for a great deal of AI infrastructure spending. The flip side: if the result fails refereeing, that thesis takes the hit with it.
For enterprise practitioners, there's something usable right now. The patterns that worked here — control objects, blind review with assigned failure modes, independent re-derivation — transfer to any verification problem, not just mathematics. Audit logic, financial model review, security analysis. "Ask one model and ship it" is measurably worse than "assign several mutually blind instances different ways to fail."
For general readers, treat the headline carefully. "AI is 67% of the way to solving the Riemann hypothesis" is wrong. It's not a progress bar, and it says nothing about where the other 32.75% of zeros sit. The Riemann hypothesis is exactly as open as it was two weeks ago, and the Clay Institute's million dollars is still on the table.
For math students, there's a sharper reading. This argument worked because nobody had tried that combination. Montgomery's pair correlation and Sylvester's law of inertia are both textbook material — they just live in different rooms, and few people hold both at once. If cross-field connection is precisely what language models are good at, the durable human edge shifts away from computation and toward deciding which question is worth asking.
🥄 Three Things You're Probably Wondering
— So is the Riemann hypothesis solved? No, and not close. Anthropic wrote plainly that they "don't expect that the techniques Claude used will lead to proving the Riemann hypothesis." This is a record in a separate problem — how many zeros are provably on the line — and 67.25% tells you nothing about the remaining 32.75%.
— Is it actually correct? Who checked?
It hasn't been peer reviewed. Two Anthropic mathematicians, Levent Alpöge and Ralph Furman, studied it, and Brian Conrey and Daniel Goldston read the manuscript on short notice — none of which is a referee report. What you do have is a public, sorry-free Lean 4 formalization of Theorems A–E, so anyone can build it and confirm there are no logical gaps. Confirming that the formalized statements say what the paper says it says is still a human job.
— Can I try this model? Not yet. It's an unreleased research version, and neither weights nor checkpoints are public, so nobody outside Anthropic can reproduce the run exactly. What is public is Appendix C.6, the paper's own methodology notes for anyone attempting something similar — you can copy the structure with a model you can actually access.
Sources
- Claude and the Riemann zeta function (Anthropic Research, 2026-08-10) — the official announcement, with the two sessions / 31M output tokens / ~60 subagents / 2,400 shell commands figures and the explicit caveat that these techniques aren't expected to lead to a proof of RH.
- More than two thirds of the zeros of the Riemann zeta function lie on the critical line (paper PDF, author: Claude, 2026-08-10) — Theorems A–E, the full record history from Hardy 1914 to PRZZ 2020, the optimized constant 0.6725, and Appendix C's account of the discovery.
- anthropics/zeta-23-lean (Lean 4 formalization) — Theorems A–E with no
sorry, Lean v4.33.0-rc2, pinned Mathlib commit, and an axiom check you can run yourself that returns only the three standard axioms. - 67% of the zeroes are on the line (informal proof note PDF) — the short version. On-line zeros give a (1,0) piece, off-line pairs give a (1,1) hyperbolic plane, and the rank inequality does the rest.
- How the two-thirds argument was found: two agent runs and their literature (PDF) — the campaign from the orchestrator's seat: sixty launches, the ledger of failures, and the "my prior is that it's wrong" exchange.
- How the one-half result was found — full transcript of subagent E2 (PDF) — 21:29 UTC Aug 3 to 01:03 Aug 4, 3h 34m, 54 messages, 53 tool calls, zero network calls. The log of the moment the discovery happened.
- An unconditional Montgomery theorem for pair correlation of zeros of the Riemann zeta-function (arXiv:2306.04799) — Baluyot, Goldston, Suriajaya and Turnage-Butterbaugh, 2023, published in Acta Arithmetica 2024. The paper that made the prime side's unconditionality explicit.
- Pair correlation of zeros of the Riemann zeta function I (arXiv:2501.14545) — the follow-up that asked what would happen if RH could be removed from Montgomery's proof. Claude's paper is a direct answer to that question.
Numbers are as of announcement and may change.



