Confidence Without Cause: Predicting and Localizing Physics-Shortcut Sensitivity in Language Model Hidden States
Abstract
A language model's confidence on a physics problem deserves trust only if it ignores details that play no part in the physics, yet models often shift their confidence when such a detail changes, for instance the temperature of a wall or the torque of a mechanically isolated second system. Unmeasured, this sensitivity undermines any use of confidence as a reliability signal. Measuring it has been hard: perturbed benchmarks cannot certify that an edit leaves the answer unchanged, and it is expensive to observe, since the effect surfaces only across several queries. We introduce DECOY (Distractor-Embedded Certified cOunterfactuals for phYsics): 155 families of four physics prompts that differ in a single irrelevant detail. A symbolic solver certifies for every family that the detail cannot affect the answer, and each family is labelled by the largest drop in the model's log-probability of the certified answer, at a threshold fixed before any probe was trained. Across two open-weight 7–8B instruction-tuned models, Qwen2.5-7B-Instruct is sensitive on 50.0% of the 140 of 155 families that yield a label and Llama-3.1-8B-Instruct on 22.9% of the same families. The largest effects occur when a distractor lies numerically close to the answer. A linear probe on Qwen's layer-18 hidden state, averaged over a family's four renderings, predicts this sensitivity before any token is generated at AUROC 0.760 (a ranking score from 0.5, chance, to 1.0, perfect) with whole families held out. Scored on a single rendering instead, a separately tuned probe reaches 0.747. A layer-16 probe frozen after training reaches 0.820 pooled AUROC on the 15 new families built afterwards, though within-cue-type AUROC is lower (0.667–0.6875, Section 5.2). Under a paired bootstrap over families the probe performs on par with a fitted token-entropy baseline (0.727) and reaches the lowest Brier score (a calibration measure, lower is better) of any predictor we test (0.204 against 0.214). Knocking out four attention heads at layer 16 cancels the sensitivity, localising it to a narrow, targetable mechanism. In Qwen2.5-7B-Instruct, a confident answer can therefore silently track task-irrelevant detail, and the failure is detectable before generation and traced to a location a defender can act on. DECOY supplies the certified ground truth and the fitted, paired-tested baselines against which future hidden-state monitors should be judged.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.