RoboCompiler: Agentic Experience Compilation for Robot Post-Training
Abstract
Successful demonstrations provide correct behavior, but their value for post-training depends on what a robot policy can already do. Long, reliable motions can dominate replay, leaving brief, difficult interactions with little supervision. We study how an agent can use the policy's execution experience to organize further learning. We present RoboCompiler, a framework for agentic experience compilation. A multimodal agent compares policy successes and failures with expert demonstrations, relates them through observable task phases, and maps diagnosed difficulties to relevant expert behavior. It then writes a supervision program that concentrates training around these interactions while preserving the surrounding skill. The program determines both how much exposure each phase receives and which observations provide that supervision. Across ten RoboTwin 2.0 tasks, RoboCompiler improves π₀.₅ mean success by 16.3 percentage points over full-demonstration training under the same learner-update budget. Experiments with LingBot-VLA and three physical robot tasks support its use across learners and in real manipulation. Component studies show the value of policy evidence and phase allocation, with a further, task-dependent benefit from local selection. These findings demonstrate a practical route from an agent's interpretation of robot experience to improved policy learning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.