See Once, Query Many: 3D Scene Affordance Segmentation via Agentic Memory
Abstract
Scene-level affordance segmentation is fundamental for embodied agents to perceive interaction possibilities in complex 3D environments. Existing MLLM-based methods face two key limitations: repetitive visual processing causes substantial redundancy, while insufficient modeling of remote functional dependencies and complex spatial relations limits their ability to ground intricate instructions. We present MemAff, a training-free agentic framework that converts a single scan into an evolving Agentic Scene Memory. Through incremental editing and on-demand visual recall, a vision-language agent anchors objects to frames and cross-links spatially separated functional pairs. For each instruction, MemAff traverses this memory to select the view best exposing the actionable region, followed by temporal propagation and cross-view 3D lifting. This agentic memory construction and retrieval enables “See Once, Query Many” without repeated visual processing. Experiments on SceneFun3D and our 3,149-instruction FunEnhance benchmark show state-of-the-art performance, reaching 30.5% AP50 and 24.9% mIoU while reducing per-instruction latency by over 20×.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.