acceptodds
Under review as a conference paper at ICLR 2027

Beyond the Last Layer: Multi-Layer Representation Fusion for Visual Tokenization

Abstract

Representation autoencoders that reuse frozen pretrained vision encoders as visual tokenizers have achieved strong reconstruction and generation quality. However, using only the last encoder layer leaves intermediate features unavailable to the decoder, limiting the reconstruction of fine visual details. We improve detail reconstruction by supplying the decoder with fused intermediate-layer features. We propose DRoRAE(Depth-Routed Representation AutoEncoder), which uses energy-constrained routing for adaptive multi-layer aggregation and incremental correction to retain a contribution from the original representation. Starting from a pretrained base model, we first train fusion with the decoder frozen, then fine-tune the decoder with fusion fixed. On ImageNet-256, DRoRAE reduces rFID from 0.570 to 0.290 and improves class-conditional generation FID from 1.740 to 1.650 (with AutoGuidance). For text-to-image generation, DRoRAE achieves a overall GenEval score of 0.600, compared with 0.560 for the base model. With encoder features and all other settings fixed, increasing fusion expert capacity improves reconstruction. Over the measured range, rFID decreases approximately log-linearly with fusion expert capacity ().

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.