Hard or Stuck? Viability-Conditioned Steering for Budgeted Reasoning
Abstract
Practical deployments of reasoning models impose token budgets to bound inference cost, but hard limits can truncate trajectories before they reach an answer. Existing control mechanisms rely on fixed schedules or predicted reasoning lengths, frequently failing to distinguish genuinely difficult questions from off-track reasoning. Consequently, they may unnecessarily intervene in trajectories that would otherwise complete naturally. This motivates trajectory-aware intervention adapting timing and strength to individual rollouts. We show that a probe over intermediate hidden states predicts whether a trajectory will finish within budget well before the limit is reached, beyond what generation position alone reveals. Its early prediction is strongly associated with question difficulty and baseline cap propensity, while subsequent increases in predicted cap risk provide a trajectory-relative signal of worsening completion prospects. Motivated by this observation, we introduce Viability-Conditioned Steering (VCS). VCS applies sparse hidden-state updates whose strength adapts online to the increase in cap risk relative to each trajectory’s initial level and to deadline pressure as the remaining budget shrinks. The update direction is estimated offline from matched hidden states of finished and capped rollouts. Across diverse reasoning models and benchmarks, VCS outperforms uncontrolled decoding in all 16 settings, boosting accuracy by 3.8–21.6 percentage points, while cutting token consumption by 13.6–34.8% and lowering cap rates by 68.0–92.6%.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.