acceptodds
Under review as a conference paper at ICLR 2027

Action-Aware Spatial Supervision for Weakly Supervised Affordance Grounding

Abstract

Weakly supervised affordance grounding (WSAG) aims to localize action-relevant object regions using only image-level supervision. Existing methods exploit exocentric human–object interaction images to support affordance localization in egocentric images, with recent approaches further leveraging vision–language foundation models to generate dense pseudo labels for the egocentric view. However, exocentric cues are still aggregated over object regions without explicit guidance from the queried action, while egocentric pseudo labels and predictions may remain spatially imprecise. We therefore propose a unified framework that improves spatial reliability across exocentric transfer, pseudo-label learning, and prediction optimization. First, we introduce an Action-Aware Exocentric Pooling (AAEP) module that combines queried-action-guided feature aggregation with fine-grained exocentric interaction pseudo-mask supervision, emphasizing action-relevant interaction regions. Second, we develop a Reliability-Calibrated Pseudo-Label Refinement strategy that performs action-aware refinement of egocentric pseudo labels and adaptively fuses the refined labels with the originals. Finally, we introduce a Peak-Anchored Support Expansion (PASE) Loss that encourages compact, appearance-consistent support around reliable peaks, reducing overly concentrated responses. Experiments on AGD20K demonstrate that our model achieves state-of-the-art performance, with a 4.2% reduction in KLD and gains of 1.6% in SIM and 4.1% in NSS on the Unseen split.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.