CONGRATTS : Convergence-Guided Resource-Aware Test-Time Scaling for Local LLMs
Abstract
Test-time scaling (TTS) can improve language-model outputs by repeated output sampling or refinement, but the number of scaling rounds required varies across inputs: some outputs stabilize after only a few rounds, while others continue to change, and the cost of another round depends on system resources currently available. We introduce CONGRATTS (**CON**vergence-**G**uided **R**esource-**A**ware **TTS**), a framework for adaptive TTS of locally-hosted models. CONGRATTS decides after each scaling round whether to continue by jointly considering output convergence and live resource feasibility. The same framework supports both self-consistency and self-refinement scaling, without retraining, model modification, hidden-state access, or an external verifier model, and is evaluated across multiple small language models, related datasets, and distinct inference backends. In our evaluation, CONGRATTS reduces self-consistency scaling rounds by 48.2% with only a 0.39 percentage-point accuracy decrease relative to fixed-budget scaling, and self-refinement rounds by 88.3% while improving accuracy by 1.96 percentage points. Under controlled resource contention on both a MacBook Pro (M1 Max) and Jetson AGX Orin, it reduces latency by 34.4-87.6% and energy use by 31.6-87.9% relative to fixed-budget scaling, with only a 0.38 percentage-point average accuracy decrease across both platforms. Our code is available at: https://anonymous.4open.science/r/congratts-E021.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.