Limits of Surprisal-Based Self-Evaluation: Non-Identifiability of Reasoning Correctness from Single-Model Output Probabilities
Abstract
Can a language model's token-level surprisal detect reasoning errors? We study selected-answer surprisal (the Information Gap) as a self-evaluation signal on GSM8K and SVAMP with Qwen2.5-3B-Instruct. Three findings emerge. First, cross-question classification yields near-chance accuracy (AUROC 0.5 on both datasets), with question-level heterogeneity attenuating a weak within-question signal. Second, a 14-condition counterfactual experiment reveals that Gap is insensitive to reasoning validity—four qualitatively distinct reasoning-error types (arithmetic, operator-swap, deletion, and semantic contradiction) are all TOST-equivalent to correct reasoning—while Gap responds to answer identity as a graded function of surface plausibility (, ). A mixed-effects model confirms answer correctness dominates (, ) while reasoning correctness contributes nothing (, ). Cross-dataset validation on SVAMP replicates answer-exposure dominance (6.3 over reasoning). Third, trajectory-dynamic features exhibit correct directional trends but fail independent preregistered validation (0/3 features survive Holm–Bonferroni correction) and are non-significant on SVAMP. We formalize these empirical patterns through an identifiability framework: under answer-exposure dominance, reasoning correctness is not identifiable from single-model output probabilities. We conclude that scalar terminal surprisal provides limited standalone verification, while trajectory-dynamic features, though theoretically motivated, lack the cross-sample robustness required for practical deployment.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.