Identifying the Appearance–Kinematics Trade-off in Cross-Embodiment Transfer
Abstract
Cross-embodiment robot learning rests on a premise almost never stated: a difference in embodiment is a difference in capability. Vision falsifies it, and two recent studies that look closely disagree about how. Neither defines a distance on both axes, so their conclusions cannot be placed on the same plot. We argue that this is not an oversight. Splitting transfer loss into an appearance part and a kinematic part requires a robot whose silhouette is held fixed while its skeleton changes, and no such robot can be built: a real robot's outline is the outer surface of its collision body. Under any distribution over real hardware the required counterfactual has probability exactly zero, so the mediation quantities are not merely noisy but unidentified, and more robots do not help. We separate two structural causal models. In nature's, a manufacturing constraint forces appearance to follow collision geometry; in the simulator's it does not, and building a simulator is in causal terms the deletion of that one edge. We prove that the estimands defined on the first model are identified by four interventions on the second. Those interventions form a strict , and its two off-diagonal cells—the only ones that pull the two axes apart—are exactly the two no hardware can realise. We give a protocol that builds them: a procedural shell generator, material and silhouette dose families calibrated once and then frozen, and field-level assertions that cells sharing a skeleton are bit-identical in physics. The answer is reported in units a reader can hold, how many percent of link length one just-noticeable difference of appearance is worth. That number is not a constant but a function of the task's contact tolerance: appearance is worth a few percent of link length where the task tolerates centimetres and about one percent where it tolerates millimetres, with an interaction term too small to make the dichotomy ill-posed.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.