PACER: Training-Free Phase-Scheduled Memory for Long-Horizon Vision-Language-Action Manipulation
Abstract
Frozen vision-language-action (VLA) policies re-encode every observation from scratch, so across a chain of subgoals they can lose track of the object they are acting on, yet fine-tuning such a policy for every deployment costs data, compute, and validation. We show that a frozen policy's own prefix key–value (KV) state can serve as memory. PACER is a training-free framework that maintains this state across replans as working memory and a capacity-one event snapshot, coordinated by a proprioceptive phase controller and periodic refresh. On LIBERO-Long with the frozen policy, PACER raises three-seed direct success from 92.3% to 95.5% (97.3% with an optional instruction-level router), including 91/150 to 149/150 on a repeated-object task, while median inference latency rises by only 0.6%. Both memories contribute to the success gain. On five physical-robot tasks, direct deployment improves aggregate success from 60% to 86% over 250 trials per condition.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.