OpenAI has published a vast collection of 722 mathematical manuscripts , grouped into 372 families of results , produced by an unpublished frontier model of the company. The release follows a series of announcements that have both impressed — and alarmed — the mathematical community, while raising serious questions about research ethics and academic practices. According to the Advisory Group on Mathematics and Artificial Intelligence ( AGMAI ), an independent advisory group of leading mathematicians formed to responsibly communicate the results, the publication includes solutions to “hundreds” of open mathematical problems.
See also: OpenAI: ChatGPT in Siri 'Exhibits Consistent Underperformance'

The results have been expected for weeks, though OpenAI has not revealed exactly which problems the model has solved or when they will be published. In September, the company said its model had “ solved more than 100 long-standing open problems in nearly every area of mathematics .” The published manuscripts also include brief descriptions of the model’s reasoning, estimates of computational cost, and statistics on the number of problems attempted. OpenAI claims that the “ average result ” used the equivalent of three hours of ChatGPT Pro thinking .
The astonishing breadth of this release isn’t just about quantity. The papers cover fields such as pure mathematics, theoretical computer science, and mathematical physics. OpenAI said the model has been trained since August 28, 2026 , and that the pace of discoveries surprised even mathematicians working in-house. The company submitted about 4,000 problems to the internal model, then selected and organized the results it deemed important enough to publish.
OpenAI and the 722 mathematical manuscripts: What the release contains
The structure of the 722 manuscripts is important to understand: they are not 722 completely independent discoveries. The manuscripts are organized into 372 families of results, where each family may include a main theorem, alternative proofs, corollaries, companion manuscripts, and different formulations of the same underlying result. This distinction is important when evaluating productivity claims, as the number of "papers" can increase significantly when related proofs and consequences are separated into individual manuscripts.
Some results also include formal proofs written in Lean, an interactive theorem proving system. Lean can mechanically check whether a formal proof follows from its specified axioms and definitions, providing a stronger form of verification than plain text. However, formalization does not automatically guarantee that the theorem is important, well-modeled, or equivalent to the intended problem. OpenAI has warned that some manuscripts may contain issues and that the collection will be updated.
The claim to have solved the Navier–Stokes problem , one of the seven Millennium Prize Problems of the Clay Mathematics Institute , has attracted particular attention. The Navier–Stokes equations describe the motion of fluids and appear in contexts such as turbulence, aerodynamics, oceanography, and climate modeling. A reliable result would require extremely careful testing, as the difficulty of the problem lies in controlling nonlinear interactions at different scales. OpenAI 's claim should therefore be treated as a result awaiting independent validation , not as an officially recognized Millennium Prize solution .
See also: Wikimedia Foundation reports action by OpenAI agents

OpenAI, AGMAI and the controversy over mathematical results
AGMAI was announced on September 21, 2026, in a post attributed to UCLA mathematician and Fields Medalist Terence Tao . The nine-member group is hosted by the Institute for Advanced Study and describes itself as independent of OpenAI . Its role is to advise OpenAI and other frontier AI labs on the review and communication of mathematical results. AGMAI 's initial recommendations called for labs to publish results promptly, use established academic channels, and disclose the model used, prompts, and computational costs.
AGMAI also urged AI companies to “ refrain from using mathematical results as marketing tools to promote their models ,” a practice that they said causes significant harm to the mathematical community. The creation of the group suggests that the issue has moved beyond simple model evaluation to one that concerns governance of science, publication rules, reproducibility, and the potential disruption of a research community by the high-volume output of a private company. Terence Tao has warned that poorly communicated work produced by AI could create a “bulldozer effect”: a large volume of results that overloads normal review mechanisms.
OpenAI described its process for this release: “ We publish the results in a GitHub repository, with protocols for paper reviews and citations. We continue to explore other community-hosted alternatives that meet the committee’s guidelines. For future releases, we are committed to further improving the quality of papers through citations, mathematical reporting, and presentation of results. ”
In the broader competitive landscape, the launch is part of a broader race to develop AI systems that can perform advanced reasoning. Google DeepMind has worked on mathematical reasoning and theorem proving with systems like AlphaProof, while Microsoft, Anthropic , and Meta are also investing in frontier models with advanced reasoning capabilities. The central trend is the move from static benchmark scores to agentic research workflows, where an AI agent can generate conjectures, search the mathematical literature, generate candidate proofs, translate them into Lean , and prepare manuscripts.
See also: OpenAI: Security researchers fired for data leak

The mathematical community seems torn between excitement about the actual speedup and concern about how the results will be published. Supporters see potential: AI can explore technical spaces too large for a single researcher, standard proof systems can reduce human error, and researchers can spend more time choosing important questions. Critics, on the other hand, express concerns about unverified proofs, inflated numbers due to fragmentation of related work, insufficient disclosure of model details, and the burden on mathematicians who have to review a sudden deluge of results. The full impact of the results will take time to assess, as mathematicians process and digest them.
