acceptodds
Under review as a conference paper at ICLR 2027

Answer-Symmetrized Representations Form an Error Field in Language Models

Abstract

Large language models may make errors that mislead users and cause harm. Monitoring internal states is promising because model activations encode rich information about generation, but a detector may learn which answer token was produced rather than whether the answer is wrong. We define the Error Field (E-FIELD) as information about whether a response is wrong that becomes readable after suppressing information about the answer token. We introduce Answer-Frame Symmetrization (AFS) to expose this signal. AFS suppresses information about what the model answered and highlights information about whether that answer is correct. On 6,400 visual questions generated by rules across 16 reasoning types, 11 detectors, four Qwen3.5 model sizes, and both in-domain (ID) and out-of-domain (OOD) settings, AFS + Variational Information Bottleneck Probe (VIB-Probe) achieves a mean macro-averaged F1 score (Macro-F1) of 71.38%, compared with 63.52% for the strongest baseline selected separately in each setting (+7.86%↑), leading in seven of eight model–domain settings. These gains recur across four model families and extend to real-world benchmarks. Code is available in the [anonymous repository](https://anonymous.4open.science/r/afs-efield-artifact-2027-3600/) for reproducibility and further analysis.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.