One Step Generation via Condition Shifting
Abstract
One-step text-to-image generation must preserve image quality and text alignment without repeated sampling updates. Trajectory-based objectives learn the mapping from noise to data, while distribution matching provides feedback on the samples that the generator actually produces. This feedback often requires an auxiliary network to estimate the generated distribution, increasing training cost as the backbone grows. Self-adversarial methods reduce this cost by learning real and fake trajectories within one model, using separate inputs for each task. We propose **APEX**, which learns the fake trajectory through condition shifting while retaining the original positive-time domain. Under the original condition, the model learns the real trajectory; under a deterministically shifted condition, it learns from its own outputs for the same text prompt. Flow matching on generated samples trains the shifted-condition prediction to follow the current generated distribution. Evaluating both velocities at the same noisy sample and time gives an internal distribution-matching correction, without a separate discriminator or fake-score network. Under score–velocity duality, this difference estimates the score discrepancy between the reference and generated distributions. We use it to construct a detached target for the one-step prediction and combine the resulting correction with the RCGM base objective. Both tasks share the existing network parameters, and inference uses only the original condition in a single forward pass. Across 0.6B–20B backbones, APEX supports full-parameter and LoRA tuning. SANA 0.6B achieves 0.84 GenEval with 0.20-second latency at one network function evaluation (NFE), and Qwen-Image 20B with LoRA reaches 0.89 GenEval at 1 NFE. Under matched training on Z-Image 6B at 2 NFEs, APEX improves over TwinFlow on all five reported metrics, including Chinese and English text rendering by 10.03 and 14.09 points. Condition ablations favor affine shifting over the other tested constructions on GenEval, DPG-Bench, and WISE.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.