Semantic Self-Flow: Relation Learning in Self-Supervised Flow Matching
Abstract
Diffusion and flow models still struggle to capture image semantics, such as spatial relations between objects. Recently, Self-Flow showed that self-distillation across different noise levels for the student and teacher views can reduce reliance on local evidence for reconstruction and improve semantic representations. However, a random student mask does not target any particular semantic dependency; for example, it leaves most of the subject at the lower noise level, revealing its location, and hence the spatial relation, without requiring the relation condition. We introduce Semantic Self-Flow (SSF), which places stronger noise where recovery depends on the semantic constraint of interest. For spatial relations, the subject is noised together with three alternative locations around the reference object, which remains at the lower noise level, making the relation condition informative for reconstruction. On ImageNet, when the relation is flipped under the same noise, both SSF outputs follow their respective relation conditions for 32% of test prompts, compared with 2% for Self-Flow, while class-conditional FID remains nearly unchanged.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.