acceptodds
Under review as a conference paper at ICLR 2027

MemAff: Retrieving and Grounding Affordances from Visual Memory

Abstract

Affordances bridge high-level task reasoning and low-level physical action by identifying what objects enable and where an agent can interact with them. However, existing affordance prediction methods typically assume that the task-relevant object is visible in the current image. This assumption breaks for embodied agents with persistent visual experience: an instruction may refer to an object seen much earlier, requiring the agent to search its visual memory before determining where to act. We formulate this setting as Memory-Grounded Affordance Prediction and introduce MemAff-Bench, a benchmark with 3,628 tasks drawn from egocentric videos and indoor panoramas. Given an intent-style instruction and a sequence of past observations, an agent must retrieve the relevant observation and localize the target object at the pixel level. By systematically increasing the amount of visual memory, MemAff-Bench reveals that existing systems struggle to identify task-relevant observations as memory grows, even when long-context models can technically process the full sequence. To address this challenge, we further develop MemAff, a training-free framework that retrieves promising observations, visually reranks them, and grounds the target object. MemAff matches or exceeds the frame accuracy of a larger Qwen3-VL-32B single-pass model while using one to two orders of magnitude fewer prompt tokens per query. Our analysis also identifies that reliable retrieval among many visually similar observations is still a central challenge, establishing memory-grounded affordance prediction as an important direction for embodied agents that must reason and act over long-term visual experience.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.