From Fine-Grained Recognition to Fine-Grained Manipulation: Semantic Visual Primitives for Grounded Multimodal Reasoning
Abstract
Multimodal large language models (MLLMs) can often describe fine-grained visual details, such as part attributes or functional object regions, but may fail to faithfully ground these cues or use them for recognition and action. This reveals a fine-grained semantic grounding gap, where intermediate evidence remains plausible in language or localized by geometry-only primitives, yet is not structured as grounded evidence for reasoning. In this paper, we propose Semantic Visual Primitive Reasoning (SVP-R), a unified framework that turns visual primitives into structured evidence that binds geometry, semantics, and action cues. Each Semantic Visual Primitive binds *where* a part is, *what* property it exhibits, and *how* it can support an action, forming an evidence graph for fine-grained recognition and manipulation. To prevent such primitives from becoming post-hoc rationales, we introduce graph binding, a shared objective that verifies the evidence graph at the node, edge, and graph levels, instantiated as Reinforced Evidence Binding (REB) for recognition and Affordance Evidence Binding (AEB) for manipulation. Extensive experiments on fine-grained recognition and manipulation show that SVP-R consistently outperforms existing grounded reasoning and VLA methods, and interventions confirm that its decisions follow the grounded evidence. These results suggest that fine-grained multimodal reasoning requires binding visual details into structured evidence for perception and action.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.