DT-VLA: Lightweight Continual Adaptation of Dynamic-Temporal Vision-Language-Action Models for Efficient Deployment
Abstract
Today's vision-language-action (VLA) models must keep learning new skills as they are asked to perform new tasks. To date, the dominant recipe has been to fine-tune and redeploy the whole model for each new task, which is slow, memory-heavy, and can even overwrite earlier skills. We present an algorithm that introduces a dynamic and a temporal structure on top of a lightweight design. The dynamic structure brings each new task online by updating only a small task-local component while the shared backbone stays frozen. What earlier skills rely on is never touched, so forgetting caused by later updates is prevented by construction. A switch between skills moves only this small component instead of redeploying the whole model, sharply cutting deployment downtime. The temporal structure records every adaptation as it happens. At deployment it reads that record to schedule each incoming request to the right component, or to flag a request that matches none of them as a genuinely novel task. Therefore, a growing skill set is served without a supplied task identifier. Moreover, memory stays modest even after many skills accumulate, since each dynamic component is tiny. Finally, on a continual stream of held-out manipulation tasks, our algorithm raises the average success rate from 83.1% to 86.9% over the strongest published algorithm in our comparison. It also cuts the switching downtime of a 1,680-request sequence from 6.1 hours to 70.5 seconds. Code is available at https://anonymous.4open.science/r/DT-VLA/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.