acceptodds
Under review as a conference paper at ICLR 2027

Shared Spectral Coordinates for Composable Embodied Skills

Abstract

Vision-language-action (VLA) models are increasingly expected to accumulate skills across embodiments with different action spaces, without retraining for each addition. We represent each skill in the spectral coordinates of a frozen VLA backbone, so that independently learned skills can be composed without retraining the backbone, pooling their training data, or aligning them after the fact. For selected backbone projections, we freeze the pretrained singular basis and represent each skill as compact gain vectors that rescale the singular values, so all skills share one coordinate system from the start. Since alignment alone does not guarantee functional compatibility, we regularize the gains toward the pretrained configuration during learning and, after composition, recalibrate only the lightweight embodiment-specific readouts. Across autonomous driving, robot manipulation, and humanoid control, the method supports two- and three-skill composition within a single backbone using only 0.34M backbone adaptation parameters per skill, versus 1.09M for rank-1 and 8.7M for rank-8 LoRA experts, while retaining competitive offline performance and effective closed-loop control. Staying close to the pretrained configuration improves robustness to interference, and readout recalibration repairs most of the mismatch introduced by composition. Because setting all skill coefficients to zero recovers the pretrained backbone exactly, skills can be activated, combined, or removed at deployment time.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.