Experts of Exit: Risk-Controlled Halting in Looped Language Models
Abstract
Looped language models allow recurrent depth to be adjusted at inference time, but adaptive depth requires choosing a stopping policy for each deployment. Candidate policies are heterogeneous: checkpoints may include trained gates, whereas confidence, convergence, and fixed-depth rules can be added post hoc. Our cross-model, cross-task evaluation of fixed-threshold stopping rules finds no reliable default among the tested signal families: their quality-cost rankings vary with maximum recurrent depth, model scale, and task. We formulate adaptive-depth inference as risk-controlled selection among exit experts. Given a user-specified quality-loss budget, the selector chooses the least expensive candidate that passes a simultaneous validation test. We implement this formulation in Experts of Exit (EoE) with EoE-Safe, which evaluates each candidate on its own autoregressive generations and constructs paired bootstrap max- lower bounds simultaneously calibrated over the candidate pool. Across 160 held-out selection trials on Ouro, point-estimate selection violates the quality budget in 55 trials, with 23 exceeding twice the budget. EoE-Safe records 13 violations, none exceeding twice the budget, while retaining a mean recurrent-step reduction. On all 20 splits of Ouro-2.6B/MATH500, where no candidate meets the validation criterion, it returns full depth. The simultaneous bounds have bootstrap-asymptotic validity. Deployment measurements distinguish analytical depth savings from hardware speedups. In our single-GPU prototype, deterministic fixed-depth skipping achieves – decode throughput over full-step execution, while the native gate and a JSD rule achieve and relative to their corresponding full-compute baselines.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.