Continual Robot Policy Learning via Variational Neural Dynamics
Abstract
Robots deployed in the real world rarely operate under a single fixed dynamics model: wind changes, payloads vary, batteries drain, and hardware wears. Yet most learning-based controllers are trained once and deployed as if learning were complete. This prevents the robot from using deployment experience to recursively improve task performance. In this work, we propose a continual learning framework that uses real-world experience to improve robot policies under hidden and time-varying dynamics. Our method learns a condition-aware dynamics model from real state-action data by combining an analytical physics prior with a neural residual for unmodeled effects. Crucially, our continual learning occurs during real-world deployment: the robot repeatedly collects new interaction data, updates its learned dynamics model, and improves the policy, in a continual loop of about 10 s. During execution, a recurrent encoder identifies the current hidden condition from recent interactions, allowing the policy to retain and reuse previously learned adaptations when conditions recur rather than adapting from scratch. Unlike meta-learning, our method can learn dynamics variations directly during deployment rather than from predefined, frozen training conditions. Through simulation studies and real-world experiments, we show that the framework improves policy performance across different platforms under a range of unobserved disturbances. In simulation under large disturbances, our method reduces the hover and tracking errors by 65.7% and 53.3% over state-of-the-art online adaptation approaches. On real quadrotor tracking task under changing disturbances, the policy recovers from recurring disturbances in roughly 1s, about 5× faster than online residual re-fitting.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.