Evaluating Test-Time Scaling in Robot Policies and World Models
Abstract
Vision-Language-Action (VLA) models and World Models (WMs) are becoming increasingly capable tools for robot decision-making, enabling action generation and prediction of future outcomes. These capabilities make sampling additional actions and their imagined futures a natural approach to scaling computation at test time. We introduce a diagnostic benchmark of sample-based test-time scaling that asks when additional samples improve action selection, and how well WM imagined futures can help with selecting the right action. Our evaluation spans VLAs, integrated World-Action Models (WAMs), VLAs augmented with action-conditioned WMs, and GPT-6 Astra with a waypoint scaffold. We organize the core metrics into three dimensions: sampling potential, measured by what is the best achievable gain from sampled actions assuming a perfect oracle selector is available (oracle headroom); selection effectiveness, measured by selector gains over random selection on the same candidate pool; and uncertainty informativeness, measured by how well dispersion in sampled actions predicts failure. Our results show additional samples expose substantial oracle headroom, reaching nearly 50 percent for Astra with a waypoint scaffold on LIBERO, but tested selectors fail to achieve the same headroom consistently. We also diagnose what is behind this action selection failure and discuss future directions for WM research.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.