acceptodds
Under review as a conference paper at ICLR 2027

Sparse Decomposition of Cross-Attention OV Circuits via Isotropic Probing

Abstract

Understanding the semantic influence of text on image generation is essential for interpreting text-to-image diffusion models. Cross-attention provides a direct interface between text and the denoising network, but attention maps reveal where token information is routed rather than what each head writes into the visual feature stream. We focus on the output-value (OV) pathway, where the value and output projections compose into a fixed linear map from token embeddings to feature updates. Motivated by isotropic random measurements in matrix sensing, we introduce IsoProbe, which probes each head's OV operator with isotropic Gaussian inputs and fits a sparse transcoder to the resulting outputs. This data-free procedure yields an overcomplete dictionary of writing directions without requiring text or image samples. The resulting dictionaries approximate OV writes through input-dependent sparse combinations of learned directions. Prompt-based analyses identify concept-associated code subsets within shared heads. Injections, patches, and ablations link selected entries to changes in object, style, and attribute-associated content, providing a means to investigate their functional contributions.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.