Misplaced, Not Forgotten: Compressed Codes Map Unseen Attribute Combinations to Seen Ones
Abstract
Models often have to handle combinations of familiar attributes that never appeared together in training, and they often fail on them. Existing work tests whether composition succeeds and which representation geometry supports it. We study where a failing model's predictions go, in models whose heads for several attributes all read one compressed representation, a shared read-out. Such a model maps an unseen combination onto one specific seen combination, a failure we call capture. The captured combination is predictable before training from the nearest seen mean image (0.76 against 0.07 chance). The relation is causal. Deleting that combination and retraining moves the prediction far more often than retraining alone, usually onto the next-nearest seen combination, and repeated deletion walks down that ordering. This holds on rendered benchmarks and on real camera photographs. The attributes remain decodable, and a composition operator on the same code recovers held-out combinations at 0.86 where the model scores 0.02, so capture is a failure of placement. Reserving part of the compression budget for each attribute does not reduce capture at matched rate. This holds on a controlled task and on frozen DINOv2, SigLIP and CLIP encoders. Specialising the read-out yields clear paired gains in held-out composition in six of seven settings and a positive but inconclusive one on MPI3D. Decoding every attribute on seen combinations does not guarantee composition. We show where these models send unseen combinations when they fail, and how changing the training support changes that destination.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.