CompFlow: Composing Velocity Fields for Multi-Conditional Generation
Abstract
Generating samples that satisfy many conditions simultaneously can be achieved with a surprisingly simple operation: composing conditional velocity fields at inference time. We introduce CompFlow, a flow-matching framework for compositional inference without retraining, architectural changes, or specialized samplers. We show that weighted velocity residuals yield the score of a Product-of-Experts composition for positive weights, while negative weights introduce quotient factors to suppress selected concepts. Beyond this score identity, we derive the exact compositional defect, characterizing when the composed dynamics transport the intended density path. Computable from residuals already evaluated during sampling, this quantity connects transport compatibility to measurable interference between conditional contributions. On CLEVR, a single-object-conditioned model is composed at inference to simultaneously control shape, color, and position for up to five objects, achieving 99.1–86.5% per-object accuracy with 30× fewer network evaluations than the evaluated baselines. With pretrained FLUX.1[dev], CompFlow combines up to five conditions in high-resolution images and improves non-spatial compositional alignment over single-prompt conditioning on T2I-CompBench. The same operator also enables controlled subtraction of contexts, objects, attributes, and styles. Together, these results connect practical inference-time concept arithmetic with a principled characterization of conditional interactions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.