acceptodds
Under review as a conference paper at ICLR 2027

Path Compatibility: Persistent Anchors for Continual VLA Fine-Tuning

Abstract

Pretrained vision–language–action (VLA) policies provide a foundation for robots to acquire new skills throughout deployment. Continual fine-tuning, however, changes the perception-to-action path and can disrupt previously learned behavior. Distillation from a previous-task teacher uses a reference that may already have drifted from the response learned when an old task was acquired. We introduce Path Compatibility (PathComp), which preserves old responses together with the interface inputs that produced them. After learning each task, PathComp caches modal tokens together with their fused latents and action distributions. During later tasks, the current fusion module and action head reproduce these acquisition-time responses from the cached tokens, while a separate modal anchor constrains the encoders on stored observations. This construction provides a fixed downstream reference while accounting for upstream drift, without a rolling teacher. We evaluate PathComp with a CLIP–LoRA–GMM policy on LIBEROObject, LIBERO-Goal and LIBERO-Spatial, spanning changes in manipulated objects, task goals and spatial relations. Across these suites, it reduces final held-out action MSE by 21.5–22.6% relative to the strongest evaluated prior-work baseline under the shared policy and offline protocol in each suite. On LIBERO-Object, a shared-design target-time control and component ablations support the use of acquisition-time interface relations for retaining learned responses as the policy adapts. Code and checkpoints will be made public after acceptance.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.