Coupling Sampling Width and Reasoning Depth for Efficient LLM Reasoning
Abstract
Inference-time scaling improves LLM reasoning, but its accuracy gains come with a rapidly growing token budget. Existing adaptive methods optimize a single axis: width-based approaches stop sampling once answers agree, risking premature commitment to an agreed-upon hallucination, while depth-based approaches prune trajectories in isolation, truncating hard but recoverable reasoning. Treating the two axes as independent leaves the budget-quality trade-off unresolved. We propose Dual-Dimensional Consistency (DDC), which couples both axes through one shared path-quality signal computed during generation. On the width axis, a confidence-weighted Bayesian rule terminates sampling only when the leading answer holds a high-confidence absolute majority; on the depth axis, trend-aware stratified pruning distinguishes persistent degradation from transient uncertainty via a spectral analysis of the confidence trajectory, gated per query without tuned constants. The budget freed by early termination is thereby redirected to the trajectories still worth extending. Across five reasoning benchmarks and five LLMs from 1.7B to 32B, DDC cuts token consumption by over while matching or exceeding strong adaptive baselines.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.