acceptodds
Under review as a conference paper at ICLR 2027

Beyond Retention: Causal Localization of Positional Interference in Memory-Augmented Vision-Language-Action Models

Abstract

Memory-augmented vision-language-action (VLA) policies can become sensitive to long and distracting histories even when task-relevant evidence remains available. We ask whether such sensitivity arises solely from retention or also from how retained memory is positionally represented downstream. On RoboMME, we construct controlled FAR/MIDDLE/RECENT histories and retention-matched counterfactuals for a memory-augmented VLA. Across 30 episodes from three task families, moving semantically matched historical evidence changes the initial predicted action chunk, and forcing the exact same retained frames does not generally eliminate this divergence. We then localize two positional interfaces in the tested memory-as-modulation architecture: per-frame temporal encoding before memory attention and serialized memory-token key positions within RoPE-based memory attention. To isolate the latter, we permute represented memory while either assigning new serialized key positions or transporting each memory unit's original key positions under matched reduction order. In FrameSamp-Modul, position transport recovers the reference action output bitwise for all 27 primary episodes in each of three fresh process blocks, whereas assigning new serialized key positions produces displacement beyond the predeclared nuisance criterion. We further reproduce this positional-transport effect in TokenDrop-Modul, which changes the upstream perceptual-memory representation while retaining the downstream modulation interface. Finally, a prospective 120-rollout closed-loop evaluation does not confirm the hypothesized behavioral improvement from recent history or temporal-position correction. Together, these results separate memory availability, positional representation, action sensitivity, and closed-loop behavior, and show that a reproducible action-level positional mechanism does not by itself imply a confirmed task-success benefit.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.