acceptodds
Under review as a conference paper at ICLR 2027

Coupling Sampling Width and Reasoning Depth for Efficient LLM Reasoning

Abstract

Inference-time scaling improves LLM reasoning, but its accuracy gains come with a rapidly growing token budget. Existing adaptive methods optimize a single axis: width-based approaches stop sampling once answers agree, risking premature commitment to an agreed-upon hallucination, while depth-based approaches prune trajectories in isolation, truncating hard but recoverable reasoning. Treating the two axes as independent leaves the budget-quality trade-off unresolved. We propose Dual-Dimensional Consistency (DDC), which couples both axes through one shared path-quality signal computed during generation. On the width axis, a confidence-weighted Bayesian rule terminates sampling only when the leading answer holds a high-confidence absolute majority; on the depth axis, trend-aware stratified pruning distinguishes persistent degradation from transient uncertainty via a spectral analysis of the confidence trajectory, gated per query without tuned constants. The budget freed by early termination is thereby redirected to the trajectories still worth extending. Across five reasoning benchmarks and five LLMs from 1.7B to 32B, DDC cuts token consumption by over while matching or exceeding strong adaptive baselines.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.