acceptodds
Under review as a conference paper at ICLR 2027

EEMAG : Egocentric-Exocentric Multimodal Affordance Grounding

Abstract

With the rise of intelligent systems, interactions between humans and robots need to be more precise and understandable. To this end, affordance detection, linking each affordance to precise objects' regions, has become more important than ever. This bridge between affordances and regions is learned by humans through a mix of experimentations and instructions. In this paper, a novel framework is presented to based on this multimodality of affordance learning, between exocentric experiments and semantic instructions : Egocentric-Exocentric Multimodal Affordance Grounding (EEMAG). It extends the multimodal information extraction from a semantic description to the multimodal description of ongoing action. This knowledge is then transferred to an Egocentric Branch through a process of distillation and fusion. To achieve this goal, we propose a new branch to replace textual branches : the Multimodal Branch. It uses a Visual Language Model and Perceivers to extract visual and semantic information about the exocentric image. This branch effectively separates affordances and guide the Egocentric Branch towards useful visual features. Our framework is trained using the weakly supervised paradigm, using only image-level labels. Different experiments demonstrate the performance of our framework compared to state of the art affordance grounding models.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.