OpenAI posts hundreds of math results from a model it has not released
On Tuesday OpenAI published a GitHub release covering 377 findings, described elsewhere as 372 result families across 722 manuscripts, in algebra, number theory, combinatorics and theoretical computer science. Many proofs have been checked in Lean. The model is not public. The company said an average result took about three hours of compute.

San Francisco3 min read
Last updated
OpenAI put a large batch of mathematical claims on GitHub on Tuesday evening, and the count depends on which desk is counting. The New York Times reported 377 results, across algebra, number theory, theoretical computer science, logic and topology. Scientific American reported 372 results from the same unreleased model. The Verge, citing the advisory group around the release, described 722 manuscripts covering 372 result families. The overlap is the story: several hundred claimed advances, grouped so that related papers count once, released together at about 18:00 Eastern on 6 October.
The model that produced them has not been released. OpenAI has said the same of the system it credited last month with a solution to the Navier-Stokes problem, one of the Millennium Prize problems, which carry a $1 million reward attached by the Clay Mathematics Institute. Tuesday's batch does not come with that prize tag. It comes with a repository, some summaries of how the model reached an answer, and a compute note. The Times said that for 10 of the findings the company provided a summary of the model's path, and that the average result took about three hours of computing.
What has been checked, and what has not
The check that matters is Lean, a proof assistant that accepts a formal argument only if the logic type-checks. Scientific American reported that many of the results have already been verified in Lean, which makes those particular claims very likely to be correct as formal statements. A Lean check does not say the result is important. It says the proof, as formalised, has no gap of the kind the checker can see. A manuscript that has not been formalised is still a claim.
Among the items named in the early reports are a solution to the four-dimensional Kakeya conjecture, improvements on some widely used algorithms, and work aimed at the Riemann hypothesis that the company describes as progress rather than a proof. Those three sit at very different heights. A formal proof of a stated conjecture is a closed file. An improvement on an algorithm is a complexity claim that other groups can time. Progress toward Riemann is a phrase that can cover a lemma or a failed approach, and it needs the manuscript before anyone outside the company can say which.
The advisory board, and the complaint it is meant to answer
A few weeks before the release, OpenAI said it was working with an advisory board of mathematicians on how companies should tell the field about machine-produced results. The Verge identified that group as AGMAI, set up to handle communication. The board exists because the September Navier-Stokes announcement landed on a field that had no shared rule for what counts as a disclosure: a theorem, a sketch, a formal proof, or a blog post. Tuesday's repository is the first large test of the new habit. It is still a company release, not a journal issue.
The objection from working mathematicians, already visible after the September claim, is about credit and about flood. A formal proof can be checked. A few hundred of them cannot be read in a week. Scientific American's point, that the batch will take months to parse, and that some proofs may be new ideas while others are recombinations of known methods, is the practical constraint. Lean can clear the logic. It cannot tell a department which ten papers to assign.
What the compute line does and does not say
Three hours of compute for an average result is an internal accounting line, not a cost in dollars and not a comparison with a mathematician's year. It does say the company is treating these as batched jobs on a model it will not yet hand over. Outside groups cannot rerun the search. They can rerun the proof, if the formalisation is in the repository. That split, a hidden search and a public check, is the arrangement Tuesday put in place.
The number to keep straight is the family count rather than the manuscript count. 722 papers can inflate a result that has been split into cases. 372 families, or 377 findings in the Times count, is the scale of the claim. Last month the company said the same model class had resolved more than 100 long-standing problems. Tuesday is the list. How many of those families survive a referee, as opposed to a proof checker, is the figure that does not yet exist.