acceptodds
Under review as a conference paper at ICLR 2027

When Unfinished Counts as Wrong: Generation Budgets Define the Target in Reasoning Monitor Evaluation

Abstract

Reasoning monitors estimate whether an ongoing reasoning trajectory will eventually yield a correct answer, enabling early stopping, verification, and runtime oversight. Many evaluations obtain the target label by matching answers from generations produced under a fixed token budget. Under a larger-budget prediction objective, however, this procedure can assign the same negative label to completed errors and trajectories that simply ran out of budget before producing their intended final answer. Our key observation is that the generation budget is therefore part of the prediction target, not merely an implementation detail. We audit this dependence across seven open models and 19 model–dataset configurations, comparing 256-token labels with paired 2048-token operational reference labels for 2,886 trajectories while holding monitor scores fixed. Heuristic truncation reaches 92.5%, label agreement falls to 28.5%, and point-estimate model rankings change. On DeepSeek-R1-7B/GSM8K, a last-step probe scores 0.892 AUC against short-budget labels but 0.568 against the reference labels. Budget sweeps and completion controls are consistent with completion-related confounding, whereas reference-label retraining yields limited, configuration-dependent recovery. These results expose a target-validity failure mode in reasoning-monitor evaluation and motivate paired label auditing before monitor performance is interpreted as longer-horizon correctness prediction.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.