acceptodds
Under review as a conference paper at ICLR 2027

Generative Modeling Beyond Pixels with Any Pretrained Vision Transformers

Abstract

Many pretrained transformers include task heads that decode their features into structured outputs. VGGT and Depth Anything V2 (DAv2), for example, produce depth maps, while VGGT also produces camera poses and point clouds. We show that the feature space of such an encoder can be modeled directly. This turns the encoder into a text-conditioned generative model whose samples are decoded by the heads it already has. Our framework, MIRAGE, compresses encoder tokens with a hybrid VAE onto a product of spheres with log-normal radii, and trains a diffusion transformer on these latents using rectified flow that follows geodesics on each sphere and a straight path in log-radius. Because generation targets the encoder's own feature space, REPA alignment can use the encoder itself as its target, and no external representation model is needed. For long videos, a temporal compressor and per-frame flow times let a single fixed-window model either generate a clip in parallel or roll forward on clean past context, producing videos of arbitrary length. We evaluate four geometric encoders: VGGT, PAGE-4D, Depth Anything V2, DUSt3R, and provide initial audio experiments. A single model generates all outputs supported by each encoder's heads. Against raw-feature diffusion baselines at two scales, MIRAGE's latents are 16-64 smaller, remains within 1-6% in depth-FVD, lowers depth-gradient divergences across all tested encoders, and samples 2 faster with 4-13 less sampling memory beyond model weights. Its compact latent also makes rolling generation of long videos practical.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.