On October 6, 2026, OpenAI made an extraordinary claim about the future of mathematics: an unreleased artificial intelligence model had produced a vast collection of research results touching hundreds of open problems in mathematics and theoretical computer science. The collection contains 372 families of results and more than 700 manuscripts. Some address questions mathematicians have pursued for decades; others offer new approaches to major areas of research.
The scale alone is striking. But the interesting question is more demanding than whether AI can produce hundreds of mathematical papers. How many of these results are correct, how many are genuinely new, and how much of their mathematics do we actually understand? Those questions matter because a mathematical claim, a computer-checked formal proof, and a discovery assimilated into the scientific literature are related achievements, not interchangeable ones.
What exactly did OpenAI release?
OpenAI published the collection in its public mathematics repository on GitHub, accompanied by a statement describing the project. The 372 figure counts result families, rather than 372 individually peer-reviewed discoveries. A family can contain a main theorem, alternative demonstrations, related consequences or companion papers.
The company describes a process in which its internal model was presented with approximately 4,000 research problems. The overwhelming majority of the selected results reportedly followed a common procedure, with an average computational expenditure comparable to about three hours of ChatGPT Pro thinking per result. That is an estimate of computational effort using the internal model, not evidence that the same outcome can already be reproduced with publicly available ChatGPT.
The repository also includes mathematical manuscripts, supporting materials, formal proofs for a portion of the results, and selected summaries of the model’s reasoning. Importantly, OpenAI itself acknowledges that the collection is at different stages of verification and that some results without formalization may contain errors. Manuscripts may be revised as researchers examine them.
Which mathematical problems make this release remarkable?
The Kakeya conjecture provides one of the clearest examples. Imagine an extraordinarily thin needle that must be turned so it points in every possible direction. How small can a set containing a unit line segment in every direction be? In higher dimensions, versions of this deceptively simple geometric question lead into deep harmonic analysis. OpenAI’s catalogue claims results concerning the Kakeya maximal conjecture in three dimensions and the Hausdorff dimension of Kakeya sets in four dimensions. These are specific mathematical statements that specialists must evaluate; they should not be reduced to a vague headline that “AI solved geometry.”
The Riemann zeta function offers another revealing case. The Riemann hypothesis concerns the locations of particular zeros of this function and is one of mathematics’ most famous unsolved problems. The catalogue includes claims about zero-free regions, including a result presented as a quasi-Riemann hypothesis. This is potentially important work related to the Riemann hypothesis, but it is not the same as proving the full Riemann hypothesis. The distinction is essential: progress on a prestigious problem should not be confused with its complete resolution.
Other manuscripts cover theoretical computer science, algorithms, analysis, algebra, probability and mathematical physics. A contribution need not settle a famous conjecture to matter. Sharper bounds, more efficient algorithms and connections between previously separate techniques can all advance a field, provided that their claims survive careful examination.
What does it mean when a proof is verified in Lean?
One particularly significant feature of the release is the use of Lean, a programming language and proof assistant that can check formal mathematical arguments. A conventional research paper presents a proof in mathematical language for other mathematicians to examine. A Lean formalization expresses a mathematical statement and its proof through precise rules that a computer can verify.
When Lean successfully checks a formal proof under its stated assumptions, that is powerful evidence that the formal argument follows logically. It substantially reduces the risk of certain hidden logical mistakes. But an important subtlety remains: the computer verifies the statement that was formalized, not necessarily every claim a human reader believes the original paper makes.
A formalization could target a weaker theorem, rely on assumptions that are not obvious from a headline, or translate a natural-language claim inaccurately. This concern is not hypothetical. An October 2026 research preprint specifically examines mismatches between AI-assisted Lean formalizations and the natural-language claims they were intended to represent, including in the debate over Navier–Stokes.
Therefore, the most useful questions are not simply “Was it checked in Lean?” but “What precise theorem was checked, under which assumptions, and does that theorem match the claimed mathematical discovery?” Formal verification and expert mathematical interpretation strengthen each other.
Why are mathematicians skeptical?
The discussion follows an earlier controversy surrounding OpenAI’s announcement about the Navier–Stokes equations, which describe the behavior of fluids and feature in one of the Millennium Prize Problems. That announcement drew attention not only for its mathematical ambition but also for disagreements over the exact problem addressed, the origins of the approach and standards of disclosure. The larger claim should still be treated as contested rather than universally established.
The October release raises a related problem of reproducibility. The AI model responsible for most of these results has not been publicly released. Without access to the model, complete prompts and sufficiently detailed computational records, independent researchers cannot reproduce the reported research process as easily as they might wish. This matters especially when a company says an unfamiliar system can solve sophisticated problems with surprisingly little direction.
Reporting by Scientific American on October 6 described sharply different reactions. Andrew Sutherland of MIT urged caution about claims of single-agent, single-prompt solutions that outsiders cannot reproduce. Daniel Litt of the University of Toronto emphasized the value of making mathematical answers public rather than keeping them secret. Terence Tao, meanwhile, criticized the extraordinary speed at which frontier AI laboratories were producing research claims. These positions reflect different concerns about a shared problem: how mathematical knowledge should be produced, checked and communicated.
Can mathematics move faster than mathematicians can understand it?
Traditionally, even a dramatic mathematical breakthrough is followed by a slower process. Researchers read the argument, compare it with earlier work, search for gaps, simplify the proof, generalize the method and discover which ideas can be reused elsewhere. The final paper is only one part of this process. Understanding changes what other mathematicians can do next.
A flood of machine-generated papers threatens to separate the production of proposed results from the community’s capacity to interpret them. If an AI system produces more proofs than researchers can inspect, a new bottleneck appears: selecting which results deserve attention and establishing how they fit into existing knowledge. A theorem that is formally correct but poorly explained may be far harder to incorporate into future research.
There is also a genuine opportunity. An AI system that assembles known mathematical tools in novel combinations, explores large spaces of possible arguments and supplies checkable proofs could make researchers dramatically more productive. What looks like a pedestrian combination of existing techniques may still be exactly the combination that solves a hard problem. Mathematical creativity does not require every ingredient to be unprecedented.
What matters is whether the resulting work expands understanding. This is as much a question about the philosophy of science and the organization of research as it is about artificial intelligence.
What should we watch next?
Three developments will be especially informative: independent evaluations of the most consequential claims; additional formalizations that clearly match the accompanying natural-language theorems; and greater access to the methods, prompts or models needed for reproducibility. Corrections to manuscripts should not automatically be interpreted as failure. Revision is part of serious mathematical research. What matters is whether errors are identified clearly and whether the strongest claims survive scrutiny.
OpenAI says it intends to work toward a responsible release of its model and support workshops and conferences devoted to understanding AI-generated mathematical results. Those efforts may help, but meaningful progress will ultimately be measured by the mathematics itself and by what the research community can establish independently.
The most significant possibility is not that AI will publish more papers than humans. It is that artificial intelligence might become a serious participant in mathematical discovery while proof assistants, human interpretation and scientific criticism continue to determine what enters the body of reliable knowledge.
Frequently asked questions
Did OpenAI prove 372 unsolved mathematical problems?
No such conclusion is warranted from the release alone. OpenAI published 372 families of claimed mathematical results, some with computer-checked proofs and others at earlier stages of verification. Their correctness, novelty and importance must be assessed individually.
Did OpenAI solve the Riemann hypothesis?
The collection reports results concerning zero-free regions of the Riemann zeta function. These are related to the subject, but they do not establish the full Riemann hypothesis.
Does Lean verification guarantee that a paper is correct?
Lean checks a formal statement and proof. It does not, by itself, establish that a natural-language description expresses exactly the same theorem or that all contextual claims about novelty and significance are justified.
Can the public use the model that produced these results?
Not at the time of this article. OpenAI describes it as an unreleased internal model and says it is working toward a responsible release.
Further reading and primary sources
- OpenAI’s mathematics repository, manuscripts and proof artifacts.
- OpenAI’s October 6, 2026 announcement.
- Research on the limitations of Lean verification of AI-generated natural-language proofs.
- Joseph Howlett’s October 6, 2026 reporting in Scientific American, which informed the questions examined here.
For an earlier example of AI contributing to research mathematics, see InsightArea’s analysis of the Erdős unit-distance conjecture.
Written by Costin Liculescu for InsightArea, where mathematics, artificial intelligence, computer science and scientific reasoning meet.

Comments are closed.