GAMA: COMPOSITIONAL OBJECT-CONCEPT BINDING FOR CONTROLLABLE MULTIMODAL EMBEDDINGS
Abstract
Frozen multimodal encoders can recognize objects and attributes while still entangling which concept belongs to which object. This limits counterfactual object-level editing: moving a feature direction for one object may unintentionally perturb another object or the surrounding scene. We propose GAMA, a post-hoc structured intervention interface that reconstructs frozen scene embeddings from concept-position tuples while exposing editable object slots. Its core mechanism combines reusable additive concept semantics with gated low-rank multiplicative conjunction residuals. Across controlled text and visual scenes,GAMA preserves frozen-space geometry while supporting targeted replacement and object-concept decomposition. Continuous-factor recovery and a GroundingDINO/CLIP pipeline test a practical relaxation from curated tuples to automatically extracted pseudo-tuples. Because GAMA reorganizes only information retained by a frozen encoder, its scope is compositional recombination within given structured concept spaces, not end-to-end perception or unrestricted openvocabulary recognition of unseen object identities.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.