acceptodds
Under review as a conference paper at ICLR 2027

Separate Adaptation, Unidirectional Coordination: Learning Bimanual Manipulation with a Shared VLA Backbone

Abstract

Learning bimanual manipulation from limited demonstrations requires both reusable single-arm skills and effective cross-arm coordination. Monolithic policies learn these jointly from costly bimanual data, while compositional methods reuse single-arm vision–language–action (VLA) models but leave the choice and use of exchanged information less explicitly modeled. We propose Separate Adaptation, Unidirectional Coordination (SAUC), which separates arm-specific adaptation from cross-arm interaction. Two independent low-rank adapters specialize a shared frozen single-arm VLA backbone. A lightweight State–Intent coordinator then uses the Follower's context to select and modulate the Leader's contextual state and action-readout intent before updating the Follower's intent representation. On six RoboTwin 2.0 tasks, SAUC trained with 50 clean demonstrations per task achieves absolute gains of 5.00% and 5.33% in mean success over the strongest evaluated monolithic pretrained policy under Easy and Hard, respectively. Compared with TwinVLA, it achieves absolute gains of 13.33% and 14.00% while using about as many trainable parameters. Ablations support separate adaptation and explicit coordination, and real-world experiments show higher success than TwinVLA on all three evaluated tasks. We find empirically that a shared single-arm action prior improves bimanual policy performance through tailored arm-specific adaptation and selective information exchange.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.