acceptodds
Under review as a conference paper at ICLR 2027

CAP: Commitment-Aware Penalties for Training and Deploying Reasoning Models

Abstract

Reasoning models often keep writing after they have internally settled on an answer, and much of this additional output does not change the final result. Reinforcement learning methods that penalize length address this by cutting text indiscriminately, at a cost that falls hardest on the problems that need long reasoning most. We therefore propose CAP, Commitment-Aware Penalty: a lightweight probe reads the model's internal commitment state from hidden states during a forward pass that GRPO already performs, so the penalty applies only to text generated after commitment. We train this probe separately on each of six models from five architectures. We then test the resulting reward in two settings: (1) At training, a single full-scale run on Qwen3-8B shows no detectable accuracy difference from outcome-only GRPO on four sets not used for checkpoint selection, while generating 9–28% fewer tokens per set. A conventional length penalty compresses more aggressively but trails CAP by 8.7 points on MMLU-Pro and 6.3 on AIME 2025; only the MMLU-Pro gap survives multiple-comparison correction. (2) At deployment, the same probe can instead be used on base models as a stopping rule: evaluated on every trace, it removes over half the chain-of-thought tokens on MATH-500 for under 3 accuracy points, losing fewer points than fixed-fraction truncation at similar savings and far fewer than confidence-triggered stopping. This operating point, however, does not transfer to harder benchmarks without re-tuning.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.