OpenAI posts 377 findings on open maths problems, with ten worked examples
On Tuesday OpenAI released 377 findings across algebra, number theory, logic, topology and theoretical computer science, from an unreleased model. It said an average result took about three hours of computing, and it published summaries of the method for ten of them. The release follows a disputed Navier-Stokes claim last month.

San Francisco3 min read
Last updated
OpenAI on Tuesday published 377 findings aimed at open problems in mathematics, covering algebra, number theory, mathematical logic, topology and theoretical computer science. The company said the work came from an unreleased model, that an average result took about three hours of computing, and that for ten of the findings it was releasing a summary of how the model arrived there. It also said it was working with an advisory board of mathematicians. The New York Times reported the release as a more cautious sequel to a September claim that an OpenAI system had solved a case of the Navier-Stokes equations, one of the Millennium Prize problems, a claim that angered a large part of the field.
The number 377 is a pile, not a theorem list the profession has checked. An open problem in this setting can be a conjecture in a paper, a step in a proof that has sat unfinished, or a question a model was pointed at from a public list. Until a named mathematician signs a verification, a finding is a candidate. OpenAI's decision to show its working on only ten of the 377 is the practical limit. Those ten can be read. The other 367 cannot be audited from the summary the company has chosen to withhold.
The fight last month explains the shape of this release. A claim to have settled Navier-Stokes, even in a special case, lands on a problem with a million-dollar prize and a century of partial results. Mathematicians who looked at the September argument said the system had leaned on existing human work and had not closed the estimate the problem actually requires. OpenAI has not, in the Tuesday account, withdrawn that claim. It has changed the format: many smaller findings, a stated compute time, an advisory board, and a short sample of method. That is a response to the charge of over-claiming. It is not a verification.
Three hours of computing per result is a cost figure, and it is worth keeping. It says these are not one-line completions a chat window can fake in a few seconds. It also says nothing about whether the hours were spent searching a proof space or stitching lemmas the model had memorised from papers. The advisory board is the body that could tell those two apart, if its members are named and if their notes are published. The Times account available on Tuesday did not list the members. A board that is unnamed cannot be cited as a check.
The fields are not equal in how an error hides. A number-theory identity can be tested on a machine for a wide range of cases and still fail on the case that matters. A topology argument fails in a diagram. A theoretical-computer-science result fails when a reduction does not preserve the quantity it claims to preserve. Algebra sits between them. A release that mixes all five under one count of 377 invites a reader to treat them as one grade of certainty. They are not. The ten summaries are the only place a specialist can start, and only in the subfield those ten happen to occupy.
The professional consequence is a queue, not a celebration. Journals and preprint servers already have a trickle of machine-assisted notes. A single firm dropping several hundred at once changes the size of the queue. Referees are not staffed for it. The useful test OpenAI could still run is public and narrow: name the ten problems, name the human authors whose work the proofs rely on, and invite a contradiction on a stated lemma. Until that is done, the 377 sit as company findings, timed at three hours each, beside an unresolved argument about Navier-Stokes.
The open question is whether any of the ten summaries contains a step a working mathematician did not already have. That is a question for the people in those fields, not for the count in the press release. Tuesday's release gives them a sample large enough to start, and a remainder too large to take on trust.
Continue reading
- News
A police officer is killed at La Pampa as miners' rifles are seized
Almanaque Digital DeskPuerto Maldonado
- Sports
Saudi Arabia ends a 22-year wait, beating the UAE 2-0 in the Gulf Cup final
Almanaque Digital DeskJeddah
- Finance
L'Oréal hires restructuring counsel as US talc cases rise to about 760
Almanaque Digital Desk