acceptodds
Under review as a conference paper at ICLR 2027

Read Where the Model Reads: Causal Anatomy of ViT Classification and Its Instruments

Abstract

In language models, the logit lens interprets intermediate representations by applying the final unembedding directly to the current token state. This is natural because the current token stream ultimately reaches the unembedding. However, in Vision Transformers (ViTs), the classifier reads only the CLS stream, while image information resides in the patch tokens. Despite this structural difference, direct projection of intermediate representations through the final classifier is routinely used to interpret ViTs. We show that this routine can be misleading. Across five ViT backbones, we first examine the causal anatomy of classification and find that class evidence develops within the patch stream and is written into CLS only through a few terminal layers, which we call the *read window*. Motivated by this anatomy, we introduce the read-through lens (RTLens), a logit lens variant that passes intermediate patch states through the read window before applying the final classifier. We further derive an anatomy-based attention map that exactly decomposes patch contributions to the prediction through the read window. Across all five backbones, RTLens predictions match the model's causal behavior far better than direct projection, and our map more accurately localizes the evidence the model actually uses. Together, these results support a simple principle for interpreting ViTs, *read where the model reads*.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.