acceptodds
Under review as a conference paper at ICLR 2027

Problem or Path? Controls for Evaluating LLM Reasoning Trajectories

Abstract

High pooled correctness prediction and rising continuation success do not, by themselves, identify within-attempt information or a prefix specific advantage. We examine two controls: within-problem evaluation and budget-indexed restart comparisons. On public R1-Distill Qwen-7B trajectories, a separately versioned early-window probe reanalysis yields pooled AUROC 0.761 but pair-weighted within-problem AUROC 0.515 on the same 22-problem set at four tokens. Within-problem intervals include chance at all ten anchors, while post-hoc within-targeted probes reveal heterogeneous positives, including a failure-weighted mean of 0.556 at 32 tokens. A complementary truncation study covers 178 problem–model cells from 89 MATH problems and two small models. One cell receives a frozen prefix-limited label, although additional cells exhibit prefix advantages. An extra-batch check on all initial terminal candidates finds higher continuation success in all ten in-grid Ministral-3 comparisons against nominal-budget restart references, nine of which are interpolated. Selection, small branch counts, and unresolved numerical differences from archived probe reports limit interpretation. The results motivate problem conditioned discrimination and explicit restart-budget controls, not a general absence of outcome information or prefix value.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.