acceptodds
Under review as a conference paper at ICLR 2027

MAGPIE: Memory-augmented Affordance Grounding via Point-cloud Instance Selection

Abstract

Language-guided 3D affordance grounding requires a model to segment, directly in a 3D scene, the functional element that an instruction refers to. The leading methods answer it with scale: an 8B–72B vision–language model operates inside the grounding loop, every query incurs its forward-pass cost, and the decision rests on captured RGB-D frames in which the element may be occluded by its own carrier or absent from every view. Such models are designed to turn language into an executable goal, whereas the choice among instances of the same type depends on where each candidate lies relative to the others, a 3D signal that prior work largely leaves unused. MAGPIE therefore restricts the large models to a coarse region and settles the question in 3D. Specifically, a frozen vision–language model parses the instruction and selects a rendered view, a frozen grounder delimits the candidate set to one region of that view, and no large model is consulted again. MAGPIE stacks three lightweight heads on frozen point-cloud features: a type head proposes candidate instances, a set-selector resolves which same-category instance the instruction refers to, and a refine head tightens the selection into the final mask. The heads consume no image, captured or rendered, and disambiguate visually near-identical instances from 3D structure alone. On SceneFun3D, MAGPIE leads the strongest point-cloud-only prior work on every reported metric, surpasses peer-reviewed RGB pipelines up to 72B on AP₅₀ and AP₂₅, and exceeds concurrent 30B systems on AP₅₀ with three orders of magnitude fewer trainable parameters.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.