acceptodds
Under review as a conference paper at ICLR 2027

Energy-Optimal Test-Time Scaling: Latency and Power Tradeoffs in Sequential and Parallel Scaling

Abstract

Scaling of test-time-computation via both longer model outputs for reasoning chains and repeated parallel samples from a model are frequently leveraged to improve task performance. While both sequential and parallel scaling increase task performance, the real-world costs of this additional test-time computation are poorly quantified with existing test-time scaling laws relying on efficiency proxies such as FLOPS or GPU-hours. In this work, we develop an analytic task performance model for parallel and sequential test-time scaling and conduct an empirical efficiency benchmark of efficiency costs for a wide range of test-time computing configurations. We demonstrate that energy use observes distinct scaling behaviors from latency and FLOPS-optimal scaling, with a preference towards longer reasoning chains and fewer repeated samples. Accordingly, we identify energy-optimal test-time configurations for the Qwen3 family of models and find that energy-optimal selections reduce inference energy use by 22% and 242% savings as compared to latency-optimal and FLOPS-optimal selections, respectively.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.