acceptodds
Under review as a conference paper at ICLR 2027

Latent-Space Counterfactual Verification for Robust Reasoning Agents

Abstract

A reasoning trajectory can produce the correct answer while relying on a brittle shortcut. We show that nominal accuracy substantially overestimates strict robust accuracy under semantics-preserving edits, by 21.9% on average and up to 65.4% across 25 model–task configurations and can even reverse model rankings on MATH-500. To detect such failures, we introduce GAUGE, a counterfactual verifier that tests whether model representations respond consistently to controlled input transformations with known answer-level invariances and equivariances. GAUGE learns a contrastive robustness metric and uses conformal calibration to provide statistically valid score thresholds, enabling reference-free scoring at test time. Across 23 configurations, GAUGE predicts which nominally correct outputs fail under held-out edits—a property, we show, predominantly of the prompt rather than the trajectory—with a conditional AUROC of 0.770, outperforming 14 published baselines. Its probe-free distillation reads the candidate’s own hidden state, adding +0.021 AUROC over confidence and length and matching pool self-consistency at zero extra generation. Using GAUGE for self-training selection improves strict robust accuracy while maintaining nominal accuracy, recovering 60–67% of available oracle gains.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.