QUOTA: Head-Sensitive Uncertainty for Budget-Interpretable Token Routing
Abstract
Inference-time compute scaling underpins state-of-the-art reasoning in large language models (LLMs), but fully autoregressive generation with large models carries prohibitive computational costs. While speculative decoding and collaborative inference mitigate this via small–large model collaboration, existing approaches suffer from coarse routing granularity, reliance on supervised training, or model-dependent uncertainty thresholds that lack interpretable resource semantics. To address this challenge, we introduce QUOTA, a training-free token-level routing framework that decouples where to deploy LLM compute from how much capacity to allocate. QUOTA ranks generation states using head-sensitive collision entropy and calibrates raw scores into percentiles via a prompt-only empirical cumulative distribution function (ECDF). Given a user-specified LLM token quota , states in the top- uncertainty tail receive one LLM token, after which routing is recomputed from the updated shared prefix, eliminating auxiliary training or routing labels. Evaluated across mathematical reasoning, scientific QA, and code generation tasks, QUOTA achieves superior quality-latency tradeoffs and precise budget fidelity: with the Qwen3-0.6B/32B model pair, it obtains a 48.55 macro-average pass@1, outperforming SpecReason by 2.90 percentage points with a 17-second reduction in per-question latency; on AIME25 with Qwen3-0.6B/8B, it delivers 42.50% accuracy using only 29% LLM output tokens, surpassing GlimpRouter and SpecReason under lower compute budgets. Across three model pairs and twelve operating points, QUOTA maintains a budget mean absolute error of 1.32 percentage points, offering fine-grained and auditable token-level collaborative inference for efficient high-quality LLM reasoning.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.