Policy-Soup: Generalizable Manipulation through Heterogeneous Expert Merging
Abstract
Robust visuomotor manipulation requires policies to operate across diverse domains. Domain-specialized experts provide a natural way to handle domain variation, but maintaining separate experts does not provide a unified solution for mixed or previously unseen combinations of domain variations. An important yet underexplored question is therefore how to consolidate their complementary capabilities into a single reusable policy. This becomes particularly challenging when experts differ in model capacity, as conventional parameter merging assumes structural compatibility. Thus, we propose Policy-Soup, a method for consolidating heterogeneous policy experts into a unified policy. Our key insight is that different policy components encode different degrees of transferable and specialized knowledge. Policy-Soup shares transferable information through learned fusion of expert visual features and layer-wise aggregation of attention outputs, while preserving specialization by retaining heterogeneous feed-forward networks at their original widths and selecting among them through sequence-level sparse routing. The unified policy is trained on pooled source-domain demonstrations, without requiring data for every combination of environmental variations. We evaluate Policy-Soup on RoboTwin 2.0, LIBERO, PushT, and real-world manipulation across cross-environment, cross-task, and cross-embodiment settings. Policy-Soup preserves strong source-domain performance, improves robustness to held-out and compositional shifts, and consolidates capabilities across tasks and robot embodiments.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.