Cross-Architecture Universal SAEs: A Shared Sparse Dictionary for Vision and Diffusion Transformers
Abstract
Universal Sparse Autoencoders (USAEs) learn a single sparse dictionary whose indices denote the same concept across several models, but the original USAE was evaluated only on vision transformers. We extend the USAE to a vision transformer (DINOv2) and a diffusion transformer (PixArt-). We train a TopK dictionary with a separate encoder and decoder for each model, map both models' tokens to a common patch grid, and add a per-token alignment loss that encourages the two models to activate the same indices at corresponding patches. On held-out COCO images, a code computed from either model explains variance in the other model's activations: from DINOv2 to PixArt- and in the reverse direction. At the average patch the two models share 43 of their 128 active indices, 38.7 times the rate expected under independent selection. Of the 12,288 dictionary indices, 74.7% are used by both models. Because the training objective rewards per-token agreement, the co-firing rate measures what the objective achieves, not whether the two architectures converge without it. We also show that agreement measures that pool codes over an image's patches are invariant to any permutation of one model's patches, so they cannot detect whether the models agree at corresponding locations. On our checkpoint, the pooled bag-of-features cosine is 0.9955, while the per-token Jaccard index is 0.203. We therefore report cross-model agreement per token and relative to a chance baseline.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.