acceptodds
Under review as a conference paper at ICLR 2027

When Does a Probe Identify a Representation? Semantic Counterfactual Validation in Language Models

Abstract

Decoding accuracy does not show what a probe relies on; for a given probe, no existing test shows whether it tracks the property or its presentation. We propose semantic counterfactual validation, which adds two controls to the evaluation of a probe. The first modifies the presentation of symbolic theories under an exact entailment check, which guarantees that every entailed fact is preserved; a probe of the same type, refitted on the changed presentation with the same training budget, shows whether a failure to transfer leaves the property decodable. The second crosses truth with presentation for a fixed question, using minimally edited twin theories that keep every entity and target concept at the same record positions and counts. All experiments were preregistered with pass or fail conditions except those labelled post-hoc. In type-inheritance theories, a Gemma-3-27B-it probe with AUROC 1.000 falls to 0.892 on a certified rewriting while a refit reaches 0.999, and under reordering a Qwen3-32B probe falls from 0.91 to 0.61 with refits recovering only the original order. On ProofWriter, whose questions need multi-hop reasoning, reordering causes no detectable loss. Twin theories that keep the queried fact true show, in four models, that probes and answers respond to how the records are linked in the original order; a probe trained to separate truth from these edits still fails the certified test. Interchange interventions show that the answer's representation at the last token carries the edit similarly (the two recovery curves differ by at most 0.06) whether or not it changes the truth, and that the statement-site probe direction carries none of this effect. In new theories rendered in random orders, the model answers correctly in every order only when the false alternative is absent from the theory, and there the probes keep their ranking across orders. Decoding accuracy therefore does not establish what a probe tracks, and succeeding the proposed certified test does not establish that the model uses the probe.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.