acceptodds
Under review as a conference paper at ICLR 2027

KEEP: Knowledge-Preserving Expert Distillation for Embodied Policies

Abstract

We seek to adapt a pre-trained VLA policy for novel manipulation tasks while preserving previously acquired knowledge. This is a non-trivial task, as naive fine-tuning causes catastrophic forgetting, and existing remedies, such as co-training with replayed pre-training data or merging the task experts with the generalist in the weight space, usually trade retention for acquisition. To address these challenges, we introduce KEEP, a method for novel skill acquisition in generalist VLA policies via the distillation of task experts in the flow-field space. The method proceeds by (i) fine-tuning a lightweight LoRA-based task expert, (ii) computing for every training sample a competence gate that compares the expert's and the frozen generalist's own flow-matching errors, and (iii) training a student, initialized from the generalist, to match the expert's velocity field where the expert is competent and the generalist's field everywhere else. We thoroughly evaluate KEEP on a real robot with 5,150 rollouts, covering generalist evaluations, three single-task settings (bottle pouring, plate stacking, breaker box flipping), and a two-task setting with joint and sequential curricula. Every model is scored on its task suites and on a 100-episode generalist suite against the pre-trained policy, LoRA task experts, full fine-tuning with weight merging, and co-training with replayed pre-training data followed by weight merging. KEEP matches the task experts on the novel tasks while keeping generalist performance within the noise of the pre-trained policy, whereas the strongest baseline loses close to fifteen points. In real-to-sim validated simulation, we further ablate every component of KEEP over more than 250,000 episodes, including the gate, the anchor, and the data mixture. Taken together, our results indicate that the flow field, rather than the weight space, is the more favorable place to compose a task expert with a generalist policy.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.