TTS-Route: Query-Adaptive Test-Time Scaling Graphs on the Fly
Abstract
Test-Time Scaling (TTS) improves large language models (LLMs) by allocating additional computation at inference time. However, existing work typically optimizes task-level compute allocation, such as using a single TTS graph across a task, overlooking substantial variation in per-query compute requirements. We therefore study a novel problem of query-adaptive TTS routing: given a weak-to-strong LLM pool and a target average compute budget, the goal is to adaptively select the model, scaling strategy, and realized TTS graph for each query. This is challenging because non-Euclidean TTS graphs are difficult to predict directly, informative configuration signals remain unclear, and global compute must be allocated adaptively across queries. To address these challenges, our pilot experiments reveal three insights: (1) TTS graphs should adapt to runtime response quality; (2) per-action success probabilities provide a stable learning target; and (3) low-cost probe responses and verifier feedback improve routing reliability. Based on these insights, we propose TTS-Route, combining a budget-aware TTS-Router for per-action success estimation and Lagrangian action selection with a training-free TTS-Planner that constructs TTS graphs on the fly using verifier feedback. Experiments on three datasets show that TTS-Route achieves the strongest performance–cost trade-off among strong baselines and generalizes to monetary budgets and out-of-distribution queries.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.