What Did We Learn from 20k Hours of Open-Source Robot Data?
Abstract
Robot learning is often constrained by the perceived scarcity of data, yet the community has already accumulated tens of thousands of hours of publicly available demonstrations. The key challenge is thus not collecting more data, but effectively reusing heterogeneous datasets collected across different robot platforms. We propose the camera-centric action representation as a unified interface for robot data. This representation captures the relation between vision and action while abstracting away embodiment-specific kinematic differences. Building on this idea, we present CamX, a unified corpus of 20k hours of open-source robot demonstrations from 37 sources, converted to a common camera-centric representation. To leverage CamX, we train CamUVA, a unified video-action model that uses the camera-centric action representation and identifies every camera view and action stream by its camera pose. This design allows the model to extract value from cross-embodiment datasets and generalize to unseen robot embodiments out of the box. Together, CamX and CamUVA provide a foundation for systematically studying both what matters—the design choices that enable effective scaling—and what is possible—the capabilities that emerge from scaling. Ultimately, through our empirical study, we hope to show a path toward scaling robot learning by effectively reusing the growing body of heterogeneous open-source robot data.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.