What Makes a Good Pre-training? Dissecting Pre-training Science for Cross-Embodiment Transfer
Abstract
The success of vision-language-action (VLA) models has fueled the consensus that aggregating cross-embodiment data is key to generalist policies. However, given the significant structural heterogeneity across robot embodiments, the efficacy of simple data scaling remains questionable. To address this, we systematically investigate the efficacy of cross-embodiment pre-training. Through real-world manipulation experiments and controlled comparisons on RoboTwin 2.0, we identify kinematic similarity as an important consideration for pre-training data selection. In simulation, pre-training on ARX, which is kinematically closer to Piper, improves downstream performance, whereas FR3 pre-training yields negative transfer. Furthermore, we demonstrate that scaling data diversity, dataset size, and training duration all improve performance on downstream tasks. Our findings suggest that future generalist robot learning paradigms must shift from random data collection to strategic data curation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.