Intra-Encoder Multi-View Fréchet Supervision for Visual Generation
Abstract
Fr\'echet Distance (FD) has recently been adopted as a training objective for visual generative models, providing distribution-level supervision by matching the feature statistics of real and generated images in frozen visual representation spaces. Optimizing FD in a single representation space can overfit to the statistics captured by that representation, producing a low in-space FD while substantial discrepancies under other representations and poor visual quality remain. FD-loss addresses this issue with FD-SIM, which jointly aligns real and generated distributions across multiple pretrained visual encoders. This multi-encoder supervision broadens representation coverage but introduces substantial additional encoder computation during training. To obtain multi-view Fr\'echet supervision without relying on multiple independent encoders, we propose Intra-Encoder Multi-View Fr\'echet Supervision (IEMV-FD), which constructs multiple distributional supervision views from the internal representations of a single pretrained visual encoder. Specifically, we extract representations from multiple layers and form both CLS-token and average-pooled patch-token views at each selected layer. Existing Fr\'echet-based evaluation metrics may use encoders that also participate in training supervision as evaluation judges, creating supervision-evaluation overlap and making it difficult to isolate generalization to unseen representation spaces. We therefore introduce held-out , which averages normalized Fr\'echet distance ratios across 6 evaluation encoders that are not involved in the training of either IEMV-FD or FD-SIM. On ImageNet class-conditional generation, IEMV-FD achieves lower held-out than FD-SIM across multiple models, including reductions from 10.76 to 7.37 on JiT-B and from 8.01 to 6.65 on pMF-B, while using a single supervision encoder with only about 11.5% of the supervision-encoder parameters used by FD-SIM. Our code and model weights will be made publicly available.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.