CAEM-BC: Condition-Adaptive Elastic Modal Flow with Bounded Continuation for Vision-Language-Action Policies
Abstract
Improving a pretrained vision-language-action policy with limited demonstrations requires learning useful action corrections and joining them to commands already executed. We present Condition-Adaptive Elastic Modal Flow with Bounded Continuation (CAEM-BC), an adapter that refines action chunks while keeping the pretrained policy frozen. A condition-dependent elastic operator defines a low-rank temporal basis for residual prediction, preserving the base proposal outside the correction subspace. Bounded continuation then joins the corrected chunk to executed history while respecting action limits. CAEM-BC uses fewer than 6M trainable parameters and trains for 3,000 steps on 40 demonstrations per task. Across three backbones and four LIBERO suites, it improves average success for every backbone. An additional Spatial evaluation improves SmolVLA success by 3.2 percentage points while reducing executed-arm jerk by 40.5%, with further gains in execution quality under new initial states and language rewrites. Controlled component interventions show that residual correction repairs task failures, while continuation preserves the repaired outcomes and improves execution continuity. These findings connect efficient adaptation to task completion through a learned correction space and history-aware execution.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.