acceptodds
Under review as a conference paper at ICLR 2027

STAM: Spatio-Temporal Adaptive Coarse-Fine Visual Memory for Generalist Robot Policies

Abstract

Generalist robot policies need historical visual evidence when the current observation is insufficient for action selection. However, compressing history can discard useful details, while indiscriminate history conditioning can introduce irrelevant information. We propose STAM, a plug-in spatio-temporal adaptive coarse-fine visual memory module for generalist robot policies. Its coarse-fine visual memory combines compressed historical context with directly retrievable visual details. A spatio-temporal recall map determines which image patches access memory and how far back to retrieve fine visual evidence. Layerwise coarse-fine memory fusion integrates both memory sources into the selected patches without modifying the downstream policy backbone. An action-guided three-stage training procedure learns recall decisions from action errors without manual recall-map annotations. Experiments across four simulation benchmarks and four real-world robot platforms demonstrate improved memory-dependent manipulation while maintaining strong general manipulation performance. On MIKASA-Robo, STAM improves average success by 33.2 and 28.0 percentage points over the original and GR00T N1.5 policies, respectively. Code will be available after acceptance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.