acceptodds
Under review as a conference paper at ICLR 2027

See Once, Query Many: 3D Scene Affordance Segmentation via Agentic Memory

Abstract

Scene-level affordance segmentation is fundamental for embodied agents to perceive interaction possibilities in complex 3D environments. Existing MLLM-based methods face two key limitations: repetitive visual processing causes substantial redundancy, while insufficient modeling of remote functional dependencies and complex spatial relations limits their ability to ground intricate instructions. We present MemAff, a training-free agentic framework that converts a single scan into an evolving Agentic Scene Memory. Through incremental editing and on-demand visual recall, a vision-language agent anchors objects to frames and cross-links spatially separated functional pairs. For each instruction, MemAff traverses this memory to select the view best exposing the actionable region, followed by temporal propagation and cross-view 3D lifting. This agentic memory construction and retrieval enables “See Once, Query Many” without repeated visual processing. Experiments on SceneFun3D and our 3,149-instruction FunEnhance benchmark show state-of-the-art performance, reaching 30.5% AP50 and 24.9% mIoU while reducing per-instruction latency by over 20×.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.