Uncovering Conceptual Geometry in Diffusion Models with Spatially Invariant Similarity
Abstract
How do diffusion models organize the visual world they learn to generate? We introduce Spatially Invariant Similarity (SIS), a training-free and model-agnostic method for comparing intermediate diffusion representations despite variations in spatial configuration. SIS establishes patch-wise feature correspondences and preserves the spatial evidence underlying their similarity. We further introduce Testing with SIS (TSIS), which aggregates these correspondences across generations to reveal concept-level relational structure. Across diverse pretrained diffusion models, TSIS uncovers coherent conceptual geometries: semantically related object categories exhibit shared coarse organization while retaining model-specific relationships. Extending beyond object categories, we show that color, material, and viewpoint exhibit distinct block–timestep trajectories, revealing where and when different visual factors become distinguishable during generation. Finally, generations that deviate from the reference geometry in intermediate representations consistently produce rarer and less similar final images. These results establish SIS and TSIS as tools for studying both the organization and evolution of conceptual structure within diffusion models.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.