FormalTCS: Benchmarking End-to-End Frontier Formal Theoretical Computer Science Research of Large Language Models
Abstract
Large language models (LLMs) have shown growing capabilities for automated theoretical computer science (TCS) research, while existing benchmarks remain far from realistic research settings. We introduce FormalTCS, a benchmark for evaluating LLMs on frontier, end-to-end TCS research. FormalTCS contains 143 instances drawn from papers accepted to top-tier TCS conferences in 2025-2026. For each instance, we annotate the theorem statements and proof processes in both natural and formal languages to evaluate the main bottlenecks of existing LLMs in TCS. Evaluations of leading LLMs reveal that current models remain far from reliably completing the full TCS research pipeline. In particular, we reveal that autoformalization is the sharpest bottleneck of TCS auto research, where the best model achieves only 11.5 for translating natural-language claims into formal theorem statements, compared with 28.6 Pass@8 when proving human-provided formal statements. Building on FormalTCS, we further develop an automated TCS research framework that generates, formalizes, filters, and proves new claims. Of 64 generated claims, only 6 ultimately pass expert evaluation and proof verification, indicating that beyond formalization, limited research taste remains another major barrier to autonomous TCS research.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.