Controlling Object-Induced Shortcuts for Compositional Zero-Shot Learning
Abstract
Compositional zero-shot learning requires recognizing unfamiliar combinations of familiar attributes and objects. Sparse training combinations make object identity a strong predictor of the composition label, allowing a model to fit observed associations without learning sufficient visual evidence for the attribute. Yet object information also provides the context needed to interpret attribute appearance. We propose , a framework that regulates this dual role across representation, optimization, and inference. A feedback adapter pools attribute evidence under complementary object-related supports and uses their directional residual to guide token updates. Two training objectives evaluate composition scores relative to object support, while an image-dependent score correction balances competition between object groups at inference. These controls address distinct parts of composition recognition: learning discriminative attribute evidence and balancing its competition across objects. Experiments on three standard benchmarks demonstrate strong compositional generalization, with state-of-the-art AUC and harmonic mean on MIT-States and open-world UT-Zappos. All three regulation stages consistently improve recognition across datasets while adding little inference cost.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.