Answer-Symmetrized Representations Form an Error Field in Language Models
Abstract
Large language models may make errors that mislead users and cause harm. Monitoring internal states is promising because model activations encode rich information about generation, but a detector may learn which answer token was produced rather than whether the answer is wrong. We define the Error Field (E-FIELD) as information about whether a response is wrong that becomes readable after suppressing information about the answer token. We introduce Answer-Frame Symmetrization (AFS) to expose this signal. AFS suppresses information about what the model answered and highlights information about whether that answer is correct. On 6,400 visual questions generated by rules across 16 reasoning types, 11 detectors, four Qwen3.5 model sizes, and both in-domain (ID) and out-of-domain (OOD) settings, AFS + Variational Information Bottleneck Probe (VIB-Probe) achieves a mean macro-averaged F1 score (Macro-F1) of 71.38%, compared with 63.52% for the strongest baseline selected separately in each setting (+7.86%↑), leading in seven of eight model–domain settings. These gains recur across four model families and extend to real-world benchmarks. Code is available in the [anonymous repository](https://anonymous.4open.science/r/afs-efield-artifact-2027-3600/) for reproducibility and further analysis.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.