Are We Ready for Proof Weaving? Understanding Agent Swarms through Fermat's Last Theorem
Abstract
A collection of correct proofs is not yet a proof of a common goal. We study how agent systems can build a common proof from earlier results, a process we call proof weaving. The public formalization of Fermat's Last Theorem (FLT) provides a completed reference for studying this ability. We introduce FLTBench, ten Lean reconstruction tasks drawn from this development. Each task withholds a target proof and a specified set of supporting reference proofs, while retaining the target statement and its permitted mathematical foundation. Continuing a proof also requires recognizing what earlier work has established, so we examine both reconstruction and the interpretation of work records. In our evaluation, all four evaluated systems complete the same one of the ten tasks. Case analysis shows checked cross-agent reuse, preserved unfinished constructions, and correction of a false completion report. In a separate diagnostic, an agent assesses 100 constructed work records, including normal proof reports, partial work, and conflicting claims of success. Its judgments disagree with the assigned completion status in 20 cases. These findings highlight proof weaving as a challenge for agent research. FLTBench and the diagnostic cases offer the community a concrete basis for studying how local contributions can become complete mathematical proofs.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.