Rethinking Cross-View Alignment: Interaction-Grounded Part Semantic Guidance for Weakly Supervised Affordance Grounding
Abstract
Affordance grounding aims to localize object regions that support specific actions. Since dense affordance annotation is expensive, weakly supervised affordance grounding (WSAG) has attracted increasing attention. Recent WSAG methods exploit exocentric human–object interactions as weak supervision for egocentric localization. They mainly extract visual cues from exocentric interactions and transfer them through cross-view feature or region alignment. However, large differences in appearance, viewpoint, and occlusion can make such visual transfer less reliable. Motivated by this limitation, we propose Interaction-Grounded Part Semantic Guidance (IGPSG), which explores interaction-grounded functional-part semantics as a complementary cross-view transfer signal. IGPSG uses foundation models to identify the functional part relevant to the Exo interaction and represents it with a semantic prototype. This prototype indicates which part should be transferred, while Ego visual features determine where it is localized. We further introduce Part–Verb Consensus Grounding (PVCG), which refines the transferred part query with multi-level Ego features and part–verb spatial consensus. Extensive experiments on AGD20K and HICO-IIF demonstrate consistent improvements over existing methods in both Seen and Unseen settings, supporting the effectiveness of functional-part semantics as a complementary cue for cross-view affordance grounding.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.