OpenAI has released a broad set of new mathematical results produced by one of its internal frontier models. The results are being published in a GitHub repository that includes protocols for paper revisions and citations. OpenAI developed this approach after consulting with independent advisors on best practices for sharing AI-generated mathematics with the research community.
The GitHub repository contains 722 manuscripts organized into 372 families, where each family groups related papers including principal results, companion arguments, consequences, or alternative proofs. Each family is classified by mathematical discipline. The collection includes results at different stages of verification, with many accompanied by Lean formalizations and more planned as they become available.
Alongside the manuscripts, OpenAI published abridged summaries of the model's reasoning for ten selected results.
Formal verification and transparency measures
The repository includes formalizations of many proofs in Lean, a programming language that allows mathematical proofs to be checked by computer. Lean is a proof assistant, which means it verifies that every logical step in a proof is correct according to defined rules. OpenAI plans to add more formalizations as they become available. To support scientific transparency, the company is also publishing details about how the results were obtained, including 10 summaries of the model's reasoning process. The repository contains estimates of compute costs in terms of ChatGPT Pro usage and statistics on how many problems were attempted. On average, each result required roughly three hours of ChatGPT Pro thinking time.
OpenAI says it wants this work to advance human knowledge and enable further progress in mathematics. The company has committed to funding workshops, conferences, and special programs to help researchers understand the significance of major results produced by AI. OpenAI is also working to release the model that produced these results responsibly. The company says it will continue to evaluate its internal frontier models on mathematics and other scientific fields, and will update its disclosure standards based on community feedback.