CRISP: Domain Generalization through Compositional Spatial Relations over Visual Primitives
Abstract
Domain generalization requires identifying stable representations that support reliable classification across domains. Domains may differ in low-level attributes, such as color, texture, or visual style, while preserving the same structural relationships among their underlying components. Existing methods primarily address these differences by improving the training process or aligning features across domains. However, since they leave this shared compositional structure implicit, they may overlook a more reliable source of invariance and consequently generalize less effectively to unseen domains. We propose Compositional Relational Invariance from Spatial Primitives (CRISP), an image classification framework that factors visual recognition into visual primitives and their relational composition. We represent these compositions using soft unary, binary, and ternary predicates over primitive locations and appearance, yielding differentiable measures of spatial and visual alignment that can be learned end-to-end. To learn primitives and relational structure jointly, we design an end-to-end architecture with three components: (1) a visual backbone that extracts generalized features, (2) a concept bottleneck layer that maps these features to primitive heatmaps with differentiable spatial coordinates, and (3) a structural scoring layer that evaluates candidate spatial relations among the detected primitives. Finally, we compute class probability from the joint evidence of its class-specific relational compositions and localized primitive appearance. We evaluate CRISP on five real-world image-classification datasets from the widely used DomainBed suite, covering shifts in depiction style, dataset provenance, and camera-trap location. To assess its ability to exploit compositional structure, we additionally evaluate it on CUB-DG, a fine-grained bird-recognition benchmark in which distinguishing species depends on combinations of part-level attributes. Across CUB-DG and the DomainBed benchmark suite, CRISP improves accuracy by a large margin of 7.9 percentage points on CUB-DG and 1.2 percentage points over existing DG methods on DomainBed, achieving the new state-of-the-art on both benchmarks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.