acceptodds
Under review as a conference paper at ICLR 2027

Direct Latent Editing for VFM-Based Image Generators

Abstract

Modern text-to-image generators train a flow model in a compressed latent space defined by an encoder. Recent work replaces the standard VAE encoder with a frozen Vision Foundation Model (VFM) such as DINOv3 to exploit self-supervised semantic features. However, editing for these VFM-based generators has not kept pace with their generation quality. We apply four inference-time editors built for VAE-based generators to VFM-based generators, at their published or tuned settings, and find that they rarely reach the target attribute. We propose Direct Latent Editing (DLE), an inference-time editing method for fully VFM-based generators with a frozen encoder. DLE precomputes one editing direction per concept from flow-model samples. At edit time it adds this direction to the encoded latent and decodes; no flow-model call, inversion or weight update is needed. A single scalar gives continuous, signed control. On three VFM-based generators we trained (CelebA faces, MIMIC-CXR chest X-rays and ADNI-1 brain MRI), DLE reaches targets that the tested baselines do not ( vs. at most on CelebA; vs. at most AUROC on MIMIC-CXR), at lower per-edit flow-model and decoder time than FlowEdit at its official configuration.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.