SLDBench v2: Active Scaling Law Discovery
Abstract
Scaling laws predict large-model performance from small experiments, but finding a reliable law itself requires costly experiments. SLDBench tests whether language-model agents can discover scaling laws when all experimental results are given upfront. We introduce SLDBench v2 with two extensions. First, we expand the corpus from 8 to 127 scenarios spanning language-model pretraining, downstream performance, and vision–language transfer, and select 30 challenging scenarios on which a reference power-law fit reaches R² ≤ 0.8; they represent at least 1.09 × 10²³ FLOPs of source-reported training compute. Second, we introduce active scaling law discovery: an agent chooses which experiments to run, pays each one’s recorded cost (e.g., training FLOPs), and revises its law as outcomes arrive, and is scored on the same hidden extrapolation targets as in the passive setting. Across 7 frontier agent systems, median R² drops from 0.72–0.75 (passive) to 0.24–0.56 (active), and prediction error rises 1.5–4.7×, depending on the system. Agents favor cheap experiments, spending 9–44% of the full-pool cost on average; the systems that spend least fail most often, and cheap purchases leave experimental conditions needed for extrapolation unobserved. Experimental planning thus remains a key open challenge for agents that assist model development in the prevailing AI-for-AI study.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.