OpenAI Publishes 719 Math Manuscripts, Only 42% Lean-Verified
⚡ Breaking News
TechCrunch AI
October 9, 20265 min read0

OpenAI Publishes 719 Math Manuscripts, Only 42% Lean-Verified

Back to News
❝

OpenAI released 719 math manuscripts across 372 research families, generated by an unreleased internal model at an average of 3 hours of ChatGPT Pro thinking compute per result. Only 42% of top-line results are Lean-verified, and just 10 manuscripts disclosed reasoning traces, drawing criticism from AGMAI at Princeton and Terence Tao over the absence of human understanding.

Executive Overview

OpenAI has published a repository of 719 mathematical manuscripts organized into 372 research families, generated by an unreleased internal model at an average of 3 hours of ChatGPT Pro thinking compute per result, after approximately 4,000 problems were posed to the model. However, only 42% of top-line results have undergone formal Lean verification, and just 10 manuscripts out of 719 disclosed the model's reasoning traces. The release has drawn direct criticism from the AGMAI group at Princeton University and renowned mathematician Terence Tao over the absence of human understanding and the lack of machine-readable metadata linking natural-language proofs to formal Lean code.

📊 Official Technical Data & Specifications Card

Technical AxisConfirmed Official Data
💰 Pricing & Usage CostNot officially disclosed for the internal model; usage was via ChatGPT Pro subscription (average 3 hours thinking compute per result). The internal model is not publicly available.
🌐 Platforms & Immediate AvailabilityGitHub (github.com/openai/math) — public repository containing a preprints/ folder (PDFs + sources), a Lean library, and a verification index. OpenAI is considering community hosting of the materials.
⚡ Performance & Speed Benchmarks719 manuscripts across 372 research families | Only 42% of top-line results Lean-verified | Only 10 manuscripts disclosed reasoning traces | ~4,000 problems posed to the model | 3 hours ChatGPT Pro compute per result.
🛡️ Security & Breach ResistanceNo security standards announced. Key documented risks: contradictions between natural-language proofs and Lean code (Cambridge/King's College paper) | Absence of machine-readable metadata linking NL to formal proofs as requested by AGMAI.
🧠 Context WindowNot officially disclosed for the internal model. Inference was via ChatGPT Pro thinking compute at an average of 3 hours per result.
🌍 Arabic Language & Regional SupportRepository and proofs are exclusively in English. No Arabic versions or RTL support for mathematical manuscripts. Not directly aimed at Arabic users.

Deep-Dive Features & Architecture

The official statement on GitHub revealed that OpenAI used an unreleased internal model to produce the vast majority of results, at an average of 3 hours of ChatGPT Pro thinking compute per result. The evaluations expanded after models reached saturation in existing mathematical benchmarks. Exceptions to this fixed procedure included work on the zero-free region of the Riemann zeta function and the proof of the Hodge conjecture for CM abelian varieties; the writing of the zero-free region Re(s) > 11/12 for the zeta function also underwent human editing for readability.

Out of 719 manuscripts, only 10 manuscripts disclosed brief summaries of the model's reasoning chain, covering specific families: bilinear correlations of multiplicative functions (007), irrationality exponent of π (017), Mahler conjectures (087), NP-hardness at the SDP threshold (102), subpolynomial bounds for arithmetic progressions (159), Kaplan-sky conjecture for direct finiteness in characteristic two (197), Mezard-Parisi formula for diluted spin glass (221), spontaneous magnetization in the quantum Heisenberg magnet (271), isomorphism of free group factors (287), and the 3D relativistic Vlasov-Maxwell system (362).

OpenAI confirmed it will maintain the public release history of the collection, recording corrections and revisions as new versions while keeping previous versions available. It also acknowledged that 'some unverified results may contain problems' and promised to fix them promptly, and is considering community hosting of these materials.

Benchmark & Competitive Performance

According to a TechCrunch report, OpenAI did not meet the standards of the AGMAI (Advisory Group for Mathematics and AI) hosted at the Institute for Advanced Study at Princeton University. The group, composed of mathematicians and AI researchers, had requested that AI-generated mathematical results include machine-readable metadata linking natural-language proofs to formal Lean proofs, a request that OpenAI's release largely did not fulfill. The 42% Lean verification rate stands in contrast to the rigorous formal verification standards expected in the mathematical community, where a proof is not considered complete until every logical step is machine-checked.

Terence Tao, one of the world's most prominent mathematicians, criticized the release for the absence of human understanding, stating that 'problems are solved autonomously by AI agents who do not care about the broader field after solving the initial target, and do not understand AI outputs well enough to answer questions or give lectures.' A paper from Cambridge and King's College also revealed at least two contradictions between the natural-language proof and the Lean code for a problem derived from the Navier-Stokes equations, highlighting the risks of unverified AI-generated mathematics.

Industry Impact & Enterprise Adoption

The release underscores both the promise and the current limitations of AI in advanced mathematics. While the volume of output—719 manuscripts in 372 research families—demonstrates the scalability of AI-driven mathematical discovery, the low Lean verification rate (42%) and the minimal disclosure of reasoning traces (10 out of 719) raise concerns about reliability and transparency. For enterprise and academic adoption, the lack of machine-readable metadata and the English-only repository limit immediate utility for non-English-speaking researchers and for automated verification pipelines.

OpenAI's decision to keep the repository public and to consider community hosting suggests an openness to external scrutiny and collaboration. However, the criticism from AGMAI and Terence Tao indicates that the mathematical community expects higher standards of formal verification and human-comprehensible explanations before AI-generated results can be fully trusted. The release may accelerate discussions on best practices for AI-assisted mathematical research, including the need for standardized metadata, formal verification, and human oversight.

Conclusion

OpenAI's publication of 719 mathematical manuscripts marks a significant milestone in AI's foray into advanced mathematics, but the 42% Lean verification rate and the limited disclosure of reasoning traces reveal substantial gaps in rigor and transparency. The criticism from AGMAI and Terence Tao highlights the ongoing tension between AI's generative capabilities and the mathematical community's demand for verifiable, human-understandable results. As OpenAI continues to refine its approach and engage with the community, the repository will serve as both a valuable resource and a cautionary tale for the integration of AI into formal mathematical discovery.

Media Source: TechCrunch AI | Official Company Statement: Original Source | Fact Verification & Analysis: AI Tools Oasis

Original Source:TechCrunch AIThis news was formulated based on coverage from TechCrunch AI

Frequently Asked Questions

How many mathematical manuscripts did OpenAI publish?

OpenAI published 719 mathematical manuscripts organized into 372 research families. Each family contains a main result plus supporting arguments or alternative proofs, categorized by mathematical specialty. The repository is available at github.com/openai/math.

What percentage of results in OpenAI's repository are formally verified with Lean?

Only 42% of top-line results are Lean-verified, meaning 58% of manuscripts lack formal software verification. OpenAI acknowledged that some unverified results may contain errors and promised to fix them promptly.

How much ChatGPT Pro compute did OpenAI use per mathematical result?

An average of 3 hours of ChatGPT Pro thinking compute per result, using an unreleased internal model. In total, approximately 4,000 problems were posed to the model during evaluation.

Why did Terence Tao criticize OpenAI's math release?

Terence Tao criticized the absence of human understanding, stating that 'problems are solved autonomously by AI agents who do not care about the broader field after solving the initial target, and do not understand AI outputs well enough to answer questions or give lectures.' A paper from Cambridge and King's College also revealed at least two contradictions between the natural-language proof and Lean code for a problem derived from the Navier-Stokes equations.

What are the ten mathematical results for which OpenAI published reasoning summaries?

The reasoning summaries covered: bilinear correlations of multiplicative functions (007), irrationality exponent of π (017), Mahler conjectures (087), NP-hardness at the SDP threshold (102), subpolynomial bounds for arithmetic progressions (159), Kaplan-sky conjecture (197), Mezard-Parisi formula (221), spontaneous magnetization in the quantum Heisenberg magnet (271), isomorphism of free group factors (287), and the 3D relativistic Vlasov-Maxwell system (362).

AI Tools Oasis

AI Tools Oasis Team

Bringing you the latest news and analysis in the world of Artificial Intelligence with accuracy and credibility. Follow us for all updates.

Related News