acceptodds
Under review as a conference paper at ICLR 2027

Divide-and-Remember: Recursive Action-Relevant Memory for Long-Horizon VLA Policies

Abstract

Vision-language-action (VLA) models struggle on history-dependent manipulation tasks, where the current observation alone does not determine the action, and the policy needs a memory of the history. Existing memory methods decide what to remember by design, for example, keeping frames with large pixel changes, and show inconsistent gains across tasks. We view what to remember as an optimisation problem. From the POMDP formulation of imitation learning, we show that the optimal memory maximises the conditional mutual information between the action and the memory given the current observation, the information the decision needs but the observation lacks. We propose Divide-and-Remember (D&R), a recursive memory method that learns to maximise this objective, scaling to long contexts while staying compute-light. It involves two strategies: (1) the selection over the full history is divided recursively into subproblems of top- selection over tokens, so that fixed-size, lightweight selectors learned end-to-end support an unbounded history, and the composition comes with an approximation guarantee; (2) all recursion blocks share one selector, which captures the selection rule common to every block and keeps efficient. On RoboMME, a benchmark of long-horizon manipulation tasks that require remembering when, where, what, and how to act, D&R achieves a state-of-the-art average success rate with consistent gains across all four suites under a budget of only tokens; real-robot experiments show the same gain.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.