nanoswe: Speedrunning SWE-bench
Abstract
Speedrunning has become a popular engineering challenge in machine learning: achieve the best performance within a fixed compute budget. We use it as a scientific tool to study how models acquire specialized skills and how far those skills transfer. We speedrun SWE-bench by training coding agents from scratch on demonstrations from stronger models. With 192 B200-hours of training compute, we reach 25.5% pass@1 on SWE-bench Verified, outperforming GPT-4o. Scaling laws fitted to smaller runs forecast this result to within 0.3 percentage points. At matched pass@1, our agents solve largely the same SWE-bench problems as fine-tunes of a model pre-trained on trillions of tokens, and generalize about as well to new bugs and codebases. Beyond SWE-bench, however, our speedrun's capabilities are strikingly jagged: on MMLU, they do worse than GPT-2. What, then, does pre-training buy for SWE-bench? Controlled experiments support a one-student hypothesis: pre-training helps a student imitate its teacher, but once we know how well it does so, how it was pre-trained barely matters. The hypothesis holds even when comparing a model trained on the web against a "vintage" one trained only on text from before 1931. Up to the SWE-bench scores we reach, pre-training is interchangeable with specialization data. We release nanoswe as an open project for studying specialization through speedrunning.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.