Sampling Headroom Is Not Selection Gain: A Compute-Value Audit of Test-Time Scaling for Video World Models
Abstract
Test-time scaling (TTS) succeeds when additional computation improves the output a system returns. For video world models, however, a larger candidate pool can contain better videos while leaving selected quality nearly unchanged. We introduce the Compute-Value Audit (CVA), a sequential framework that traces this gap through four requirements: available improvement, informative inference-time signals, useful interventions, and gains that exceed matched uniform compute after generation and verification costs. On 192 Physics-IQ scenes, expanding the Wan2.2-5B pool from 4 to 16 candidates raises oracle quality by +9.23 IQ points (95% confidence interval [+7.44,+11.14]), while the estimated changes for denoising-stability, reconstruction-consistency, and learned-preference selectors range from −0.75 to +0.26 IQ. On VideoPhy2, perturb-and-refine sampling yields negative estimated changes in joint semantic and physical success across three fresh-seed replications at a matched budget of 60 denoiser evaluations; the intervals include zero. Reasoning controls expose the decisive distinction between predicting difficulty and identifying where additional sampling helps. An anchor–explorer policy beats matched uniform allocation on PRM800K, with gains concentrated in a small subset of problems, while privileged real-future references improve video selection. CVA thus shifts the scaling question from how much quality a pool contains to how much a realizable decision recovers at full cost.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.