POSS: Learning from Partially Observed Object States for Compositional Zero-Shot Learning
Abstract
Compositional zero-shot learning (CZSL) requires visual–semantic knowledge that can be reused in novel state–object combinations. Yet distinguishing seen composition labels need not produce such knowledge: competition between labels can discourage state evidence that remains visually valid. We study this mismatch through the semantic scope of supervision. A composition label specifies some properties of an object, while its classification objective compares it with candidates describing other properties. We introduce Partially Observed State Supervision (POSS), which organizes existing labels into sparse semantic slot–value facts and learns object-conditioned visual evidence only from the facts each annotation reveals. Shared state predictors pool supervision across training compositions, and their evidence contributes to the ranking of both seen and unseen candidates. POSS retains standard single-composition prediction and requires no additional image-level training annotations. Experiments on three benchmarks and multiple CZSL architectures show improvements in closed- and open-world recognition. State-retention diagnostics show improved preservation of visually supported states, while transfer analyses link reusable state evidence to gains on unseen compositions.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.