Let it Settle: Allocating Test-Time Compute in Looped Language Models
Abstract
Inference compute in looped language models can be allocated to additional loops or longer chains of thought. Under a fixed budget, this allocation involves a trade-off because each additional loop increases both prompt-processing and per-token generation costs. We address this with SETTLE, an allocator based on answer settling, which separates when intermediate answers stop changing from whether they are correct. Using these measurements and a small labelled calibration set, SETTLE selects a loop count and chain-of-thought token limit under an average compute budget. We evaluate six models from the Ouro, Huginn, and Recurrent-Llama families on ten reasoning benchmarks, bound selection loss in terms of accuracy-estimation error, and isolate the contributions of estimation, budget sharing, and validation. We find that the best depth and chain length vary by model, task, and budget, and that most of the improvement comes from task-level allocation, with smaller benefits from varying settings across prompts. At half the baseline cost, SETTLE-H is points more accurate than token-only allocation where both are feasible. Including prompt processing, it reaches baseline accuracy at a median of of baseline cost, allowing a one-point tolerance in four cases. Budget-conditioned self-distillation further improves GSM8K accuracy by points at a -token budget.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.