acceptodds
Under review as a conference paper at ICLR 2027

BridgeAfford: Bridging Query–Target Gap for Generalizable 3D Affordance Grounding

Abstract

Language-guided 3D affordance grounding aims to localize functional regions on 3D objects according to interaction intents. However, existing methods exhibit poor generalization to unseen objects. One challenge is the query–target gap: a semantic discrepancy between query features encoding interaction intents and target-region features, leading to inaccurate grounding. End-to-end training favors query representations tailored to seen objects, limiting their compatibility with shifted target-region features on unseen objects. To address this issue, we propose BridgeAfford, a framework that reduces the query–target semantic gap from two complementary perspectives. First, Interaction Region Alignment (IRA) promotes cross-object alignment among geometrically similar interaction regions during training. This enables structures on unseen objects that resemble previously observed ones to produce representations similar to those learned during training, facilitating target-region recognition. Second, Visual Evidence Aggregation (VEA) leverages a 2D affordance model to refine the attention distribution and guide the query toward the correct region, yielding more target-aware queries. Extensive experiments on the LASO and PIAD datasets demonstrate that BridgeAfford outperforms existing methods, particularly on unseen objects, highlighting its robust generalization ability.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.