Plan First, Act Later: Streaming Reasoning for Executable Memory in Robot Control
Abstract
Memory-dependent manipulation requires a robot to retain task-relevant evidence beyond its short observation window. Existing memory-augmented policies typically retrieve relevant information from history and reason over it at execution time, placing a heavy reasoning burden on the high-level model. To address this challenge, we propose PAL (Plan first, Act Later), a hierarchical memory system designed to alleviate this reasoning burden. PAL maintains a short-term visual buffer, a semantic-event memory, and an executable subtask list. A high-level vision-language model (VLM) converts streaming visual observations into the subtask list. At execution time, the VLM selects the currently relevant subtask from the subtask list and passes it to a low-level vision-language-action (VLA) policy. To teach the model what to write to memory and when to update it, we introduce a two-stage procedure that first densifies sparse update supervision and then uses self-labeling with task-grounded verification to generate training data that better match the state distribution encountered during execution. We validate PAL on memory-dependent manipulation benchmarks and in real-world manipulation tasks. On RoboMME, PAL achieves a 56.05% success rate, compared with 44.51% for the strongest baseline, FrameSamp-Modul. In real-world experiments across four manipulation tasks, PAL reaches 60.8% overall success, outperforming MemER (45.0%) and FrameSamp-Modul (28.3%). These results demonstrate that proactively writing future-relevant subtasks into hierarchical memory improves long-horizon visual reasoning and robot control.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.