CalibAll: Offline Eye-to-Hand Calibration for Heterogeneous Robot Datasets
Abstract
Cross-embodiment robot learning requires action representations with consistent semantics across robot platforms. Existing representations suffer from platform-specific inconsistencies, while current solutions often maintain embodiment-specific action heads or learn latent action spaces without explicitly establishing a shared geometric reference. Recent approaches adopt camera-centric motion representations to address this issue, but missing camera extrinsic annotations limit their coverage across existing datasets. We present CalibAll, an optimization-based, robot-independent pipeline for offline camera extrinsic annotation. CalibAll initializes extrinsics via temporal Perspective-n-Point (PnP) using a tracked end-effector point, then refines the estimate by aligning rendered and observed robot masks through differentiable rendering. It additionally derives camera-frame tool center point (TCP) poses from robot and gripper URDFs to support action unification. We apply CalibAll to 16 datasets spanning four robot platforms, producing approximately 97K calibrated episodes across single-arm and bimanual systems. Downstream experiments show that using the annotated dataset for cross-embodiment pretraining improves policy performance in both simulation and real-world tasks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.