A Study of Naturally Occurring Mathematical Errors in Real-World Theoretical Research with Frontier Models
Abstract
Recent advances in artificial intelligence have enabled frontier models to play a growing role in theoretical research. AI for mathematics, in particular, has attracted widespread attention. Researchers increasingly turn to proprietary models such as Fable 5 and GPT-5.6 to investigate research-level open questions. However, these models remain prone to mathematical errors, making it important to understand how such errors arise and what they imply for the responsible use of AI in theoretical research. Existing studies often examine mathematical errors through evaluations on selected problems, particularly in competition- and Olympiad-level mathematics. Systematic documentation and analysis of mathematical errors encountered during sustained, real-world theoretical research with frontier models remain limited. In this paper, we present 65 source-linked mathematical error records drawn from frontier-model investigations of 12 research-level problems in quantum information, quantum error correction, quantum complexity, and related fields. Each record links an erroneous claim or inference to its original source, relevant context, and a checkable witness explaining the mathematical defect. Where the evidence permits, we also examine visible reasoning summaries alongside subsequent outputs. Using an established classification framework, we find that logical and reasoning errors dominate our corpus, accounting for 68.2% of error labels. Overall, our study provides an empirical analysis of mathematical errors encountered in real-world theoretical research with frontier models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.