acceptodds
Under review as a conference paper at ICLR 2027

What Can a Success-Set Flip Tell Us? Independent Response Validation in Language Model Evaluation

Abstract

We ask which claims about items selected by a finite-budget success-set flip survive responses withheld from selection, larger discovery budgets, and later checkpoints. Among 151 Llama export-comparison flips selected at eight attempts, independent responses support 81 original directions; a separate pool supports 96, favors 21 opposite directions, and leaves 34 unresolved. Nine reversals survive Fisher/Holm correction over both directions of all selected items, with a median signed effect of percentage points. Across saved request-index prefixes, corrected determinate directions rise from 56 at depth 128 to 91 at 512. Increasing the discovery budget to 64 makes 120 selected items solved at both endpoints, including 57 with independent directional support. A separate Qwen3-4B checkpoint study follows 95 items with independently confirmed improvements. Their fixed-denominator score declines by points after , with a model-conditional posterior interval of , although 60 items retain support relative to baseline. An unmatched reference supplies descriptive context. Together, these fixed-item measurements connect a selected label to its later directional evidence under frozen machine scores. Training provenance and response dependence limit broader interpretation. The anonymous supplement provides analysis code and item-level results, which we will also release publicly.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.