Keeping Memory In-Context:Noise-Matched Training for Long-Horizon Robot Control
Abstract
Long-horizon robot control requires acting on information that has left the current observation. Memory-augmented vision–language–action policies address this by storing past observations and retrieving them into the policy's context, and recent work has steadily improved what is stored and how it is retrieved. Such systems nevertheless fall far short of what the same policy achieves when the relevant memory is handed to it directly. We show that much of this gap lies not in the memory but in its use: the memory does reach the policy, but the policy was trained on memory in a form different from the one it receives at deployment, and so it does not read it. We propose , which closes the training–deployment mismatch behind this read-out gap by training the policy to read the memory it will be given at deployment. reaches 42.0 on with the same backbone as the benchmark's own memory mechanisms, matching or exceeding the per-task scores of the strongest published configuration on 7 of its 16 tasks, reaches the best published M(1) average on (83.8), and on two physical arms raises success from 25.0 to 47.5 on a retrieval task and from 36.7 to 66.7 on an occlusion task. Whether a policy uses memory is decided less by how much is stored than by how it was trained to read it.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.