Singularity-Bench: Can Agents Build Agents?
Abstract
Frontier AI agents increasingly succeed on realistic, long-horizon tasks and are beginning to improve the systems through which they operate. Yet existing benchmarks often restrict agents to narrow action spaces, use simplified environments, or rely on verification methods that are insufficient for arbitrary agent-generated systems. We introduce SINGULARITY-BENCH, an open-ended benchmark for autonomous AI development with five end-to-end tasks spanning pre-training, post-training, inference, cluster orchestration, and the agent harness. Each task requires agents to implement core systems and evaluates them against mature human-engineered baselines on identical hardware. We also develop an evaluation framework that combines isolated execution, protected logging, and independent judges to detect invalid shortcuts. Across 12 agents built from 8 frontier models and 3 harnesses, Opus 5.5 with Claude Code is the only agent to outperform the human baseline on all five tasks, achieving a singularity index of 1.275, suggesting that it might already reach singularity. The remaining agents average only 0.56, with 8 of 12 scoring zero on inference, revealing highly uneven progress. Our vision is for SINGULARITY-BENCH to track when agents reach the singularity: the point at which they transition from assisting AI development to taking end-to-end responsibility for building stronger successors.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.