MorphAct: Contextual Parameterization of Vision–Language–Action Models
Abstract
Vision–language–action (VLA) models seek to support diverse robot behaviors within a shared policy. Conventional fine-tuning, however, fits a single static update to operations whose perceptual and control demands change even within one task. When these operations favor competing parameter adjustments, fitting their mixture can leave individual operations insufficiently adapted. Our atomic-task diagnostics reveal competing update directions in both vision–language and action parameters, with source-improving updates increasing loss on other tasks. This motivates MorphAct, a contextual parameterization framework that shares an adaptation rule rather than a fixed update, allowing adaptation to specialize without splitting the policy into operation-specific experts. Two coupled generators produce context-dependent low-rank updates for the vision–language and action modules. Their coupling follows the policy's computational dependency: vision–language updates shape the representations that condition action-side parameter generation. Both generators are trained jointly through the native action objective, without operation labels or expert-adapter targets. MorphAct improves average success rates across three VLA backbones on LIBERO and Meta-World, with gains of up to 15.13%. In real-world evaluation, it achieves gains of up to 17% over π₀. Ablations support the benefits of contextual parameter generation and representation coupling.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.