SAGE: Salient Factor Discovery and Generation with Visual Foundation Representations
Abstract
Given a target dataset, such as faces with eyeglasses, and a background dataset, such as faces without, contrastive analysis separates salient factors specific to the target from common content shared by both. We aim for salient representations that capture target-specific detail in each image, such as the shape, color, and position of the glasses, so that they reveal subtypes without subtype labels and guide the generation of new examples of a discovered subtype, even one with no name or text description. We introduce SAGE, which learns both factors directly in the high-dimensional spatial latent of a frozen representation autoencoder and conditions a diffusion transformer on the learned salient representation of a reference image. On Digits-ImageNet and FFHQ eyeglasses, SAGE combines high-fidelity reconstruction (rFID below 2) with unsupervised subtype discovery, recovering the digits better than baselines (probe accuracy 0.950 vs. at most 0.281) and revealing eyewear types, finer sunglasses styles, and mislabeled images; salient-conditioned generation raises Digits-ImageNet subtype accuracy over the unfactorized latent (90.5% vs. 27.7%) and diversity on both datasets. On retinal OCT, SAGE's salient space separates three diseases using only normal/disease labels.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.