acceptodds
Under review as a conference paper at ICLR 2027

Cross-Architecture Universal SAEs: A Shared Sparse Dictionary for Vision and Diffusion Transformers

Abstract

Universal Sparse Autoencoders (USAEs) learn a single sparse dictionary whose indices denote the same concept across several models, but the original USAE was evaluated only on vision transformers. We extend the USAE to a vision transformer (DINOv2) and a diffusion transformer (PixArt-). We train a TopK dictionary with a separate encoder and decoder for each model, map both models' tokens to a common patch grid, and add a per-token alignment loss that encourages the two models to activate the same indices at corresponding patches. On held-out COCO images, a code computed from either model explains variance in the other model's activations: from DINOv2 to PixArt- and in the reverse direction. At the average patch the two models share 43 of their 128 active indices, 38.7 times the rate expected under independent selection. Of the 12,288 dictionary indices, 74.7% are used by both models. Because the training objective rewards per-token agreement, the co-firing rate measures what the objective achieves, not whether the two architectures converge without it. We also show that agreement measures that pool codes over an image's patches are invariant to any permutation of one model's patches, so they cannot detect whether the models agree at corresponding locations. On our checkpoint, the pooled bag-of-features cosine is 0.9955, while the per-token Jaccard index is 0.203. We therefore report cross-model agreement per token and relative to a chance baseline.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.