OpenAI’s release of hundreds of claimed solutions to difficult open mathematics problems is forcing the field to confront a basic question: Is a proof genuinely useful when a machine can verify it but people do not yet understand it? TechCrunch reported on October 8 that the company published 719 manuscripts while falling short of several standards recently proposed by a group of leading mathematicians.

The Advisory Group on Mathematics and Artificial Intelligence, hosted by Princeton’s Institute for Advanced Study, issued guidelines in late September for frontier AI labs pursuing open problems. Its first recommendation was that labs stop using proprietary models to test advanced mathematical questions. OpenAI’s new release openly describes an evaluation of proprietary models on unresolved research problems, putting the project in direct tension with that request.

A mathematical proof and its formal code version diverge at several highlighted connections.
Researchers found discrepancies between a natural-language proof and its Lean formalization, raising questions about automated translation.

OpenAI did follow some of the group’s recommendations, including releasing results quickly and describing parts of the process used to obtain them. The transparency was uneven, however. TechCrunch found that only 10 of the 719 manuscripts included the model’s chain of thought, limiting researchers’ ability to inspect how most of the conclusions were reached.

Formalization presents another complication. AI systems can produce a natural-language argument and then translate it into Lean, a programming language that checks whether a formal proof compiles. Formal verification can establish that the encoded argument is internally valid, but it cannot by itself guarantee that the Lean code faithfully represents the explanation written for human readers.

That distinction became concrete in a paper by mathematicians from the University of Cambridge and King’s College London. The researchers documented at least two discrepancies between the natural-language argument and the Lean formalization behind an OpenAI solution related to the Navier–Stokes equations. The discrepancies do not necessarily invalidate either version, TechCrunch reported, but they raise doubts about relying on models to translate and certify their own work without independent scrutiny.

A seminar table prepared for review sits inside an archive filled with sealed AI-generated proof folders.
Mathematicians say publication is only the start: claimed results still require explanation, scrutiny, and integration into the field.

The advisory group had recommended machine-readable metadata connecting natural-language and formal artifacts. OpenAI did not include that correlation in the release. TechCrunch also reported that 42 percent of the proofs had not undergone formalization, leaving a substantial portion of the manuscripts without that additional layer of machine checking.

The deeper issue is responsibility after publication. Human mathematicians normally defend new results, answer questions, present their reasoning, and help the community identify reusable ideas. Terence Tao argued that AI-prompted solutions can arrive without anyone prepared to perform those duties. Harvard professor Melanie Wood told TechCrunch that when manuscripts are released without human understanding, publication marks the beginning of the work rather than its completion.

OpenAI’s output may still contain important discoveries, and the reported translation problems do not prove that the underlying mathematical claims are false. But volume is not the same as accepted knowledge. Until independent mathematicians can connect the formal code, natural-language arguments, prior literature, and broader implications, the 719 manuscripts remain a large research agenda for humans—not a settled catalog of solved problems.