LATTICE: Latent Imagination for Compositional Embeddings
Abstract
Deep networks often fail on new combinations of familiar concepts, such as a blue triangle after seeing blue objects and triangles only separately, even in simple synthetic settings. Handling such combinations, an ability known as compositional generalization, is a hallmark of human intelligence and essential because the space of combinations grows exponentially while data does not. Most prior work pursues it through disentanglement, which attempts to isolate the attributes, yet evidence on whether disentanglement helps is mixed and on its own it has not solved the problem. We therefore propose to also train models explicitly to compose attributes. We introduce LATenT Imagination for Compositional Embeddings (LATTICE), a training procedure that makes composition a task through analogy: just as king man woman queen in word embeddings, a blue triangle should be obtained from a blue circle by applying the change that turns a red circle into a red triangle. Given three training images that form a attribute grid with a fourth, the model must imagine the representation of the fourth by this vector analogy. On six compositional generalization benchmarks, LATTICE raises the average accuracy of lightweight backbones by a relative – without adding parameters, letting one surpass a – larger backbone designed for this problem. It also makes representations measurably more parallel, even on unseen combinations: changing an attribute moves the representation in the same direction regardless of the others, a geometry also observed in the brain. Backbones that already provide this geometry gain little from LATTICE. Ablations show that each of its components is needed. We complement these results with a theoretical analysis.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.