Identity-Preserving Attention Readouts for Deductive Validity
Abstract
Whether a conclusion follows from its premises can be read from where the conclusion's own attention goes, nearly as well as from the hidden states that a supervised probe reads. We train a logistic classifier, a readout, on a few attention values per head, such as the share of the conclusion's attention that lands on the premises. It reaches at least 89% of the probe's ProofWriter AUC in every one of the thirteen models where the readout is extracted, up to 27 billion parameters, without access to the residual stream. On Qwen3-1.7B it scores above LapEigvals, a trained detector of attention Laplacian features, on every benchmark we test, while reading fewer features per head. The reason is that statistics which stay the same when the tokens are relabeled, such as attention Laplacian spectra, cannot represent which token attends to which. We call this limit permutation blindness and state it as an elementary lemma. Spectra of the residual stream, the very states the probe reads, also stay near chance, so the limit lies in the summary and not in the tensor it reads. Even the entropy and the maximum of the conclusion's attention row, which ignore which token receives which share, beat every spectral summary of the whole attention matrix in each of those models, so detectors built on such summaries should read that row instead. The classifier also attributes the verdict to single heads, at about the per-item storage of a probe at one fixed layer, and needs more training labels than the probe. We also find that a single negation feature reads the labels of ProofWriter far above chance, invisibly to the usual length control, and release rebalanced splits and the code that regenerates every number. Code: https://anonymous.4open.science/r/latent-validity-8074/
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.