EVA: A Memory-Augmented Harness for Long-Horizon Embodied Tasks
Abstract
Recent advances in Vision-Language-Action (VLA) models have improved the generalization of embodied foundation models, yet their performance remains limited on tasks requiring long-horizon reasoning, persistent memory, and robust execution under environmental noise. Moreover, consolidating diverse capabilities into a single model is costly and difficult to scale. Inspired by general-purpose agent harnesses, recent works have explored agentic frameworks for embodied tasks, but these approaches largely rely on directly adapting general-purpose harnesses, leaving the design of a dedicated, efficient, and general-purpose embodied harness largely unexplored. Therefore, we introduce **EVA**, a general and extensible embodied harness that provides a lightweight set of essential components, integrating perception, memory, memory-grounded planning, execution, monitoring, and verification. At each step, EVA generates an executable subtask and selects an appropriate executor, monitors its execution online, verifies the outcome, and updates memory to enable replanning or recovery. This decoupled design enables existing embodied models to achieve robust, persistent, and recoverable long-horizon behaviors without modifying their parameters or requiring test-time adaptation. We further theoretically characterize the benefits of the harness-based approach and conduct extensive experiments across distribution shifts, long-horizon reasoning, and embodied-memory benchmarks. EVA achieves an average improvement of **37.8%** over strong baselines while reducing token overhead by **10.9** on VLABench. Real-world deployment on a humanoid robot further demonstrates its effectiveness. Overall, these results show that \ours can substantially extend the capabilities and generalization of existing embodied models. Code is available at [https://anonymous.4open.science/r/eva-0](https://anonymous.4open.science/r/eva-0).
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.