What You Measure, Ablate, and Pool Decides the Conclusion: Auditing Sensorimotor Pretraining
Abstract
Statistical rigour does not protect against instrument artefacts. We report a preregistered multi-body audit of a developmental sensorimotor curriculum — four MuJoCo MJX embodiments (DM humanoid, Unitree G1, Go2 quadruped, Franka Panda), a 318-run grid plus follow-ups, Holm families, bootstrap CIs, TOST — in which three routine choices each independently changed what we concluded: which *metric* you read, how you *ablate* a component, and whether you *pool* across bodies. The first indicts our own instrument. The curriculum's headline training-time advantage (Hedges up to , raw ratios up to ) survived preregistration, four embodiments and Holm correction, yet a gain-matched diagnostic shows it is largely manufactured by the curriculum's own action-gain schedule. A fixed evaluation gain amplifies the curriculum arm's measured motion by – and the control arm's not at all; removing that asymmetry removes – of the raw ratio (– of the log-ratio). Measured at each policy's own training gain on a corrected trainer (six instrument fixes), the curriculum improves the proxy on **0/4** bodies. Nor does it transfer, at final performance: on **none of four** downstream routes does the curriculum improve on its own ablation, and on the frozen routes the audited intrinsic-motivation objectives (curriculum, empowerment, prediction-error; ICM/RND) sit at or near *random init*, while a self-supervised dynamics control, itself reward-free, transfers near-ceiling and none overtakes random init at a pretraining budget (P27; ICM health caveat in Appendix K). Two further lessons reach beyond this system, on less direct evidence. First, ablating prediction error by *reward weight* understates it: removing the *component* shows it suppresses motion (); under one matched evaluation the gating gap on H1_dm is larger by units, and on Go2 the raw difference is borderline; the representation effect reproduces on shared-encoder ICM and, dose-dependently, RND. Second, a cross-body diagnostic that looks predictive pooled () is a leave-one-body-out Simpson's paradox (every diagnostic's mean , best ). That pooling failure reproduces on *released benchmark data*: naive pooling erases of a real within-task Atari advantage, and under URLB's own seed noise no "best" unsupervised-RL method can be crowned. Running that benchmark's standard pipeline ourselves ( finetunings), no method tops all three of its domains at either seed budget we ran; the sharper claim — that two poolings of the same runs disagree about the winner — held at three seeds and *failed* at five, and we print that failure rather than the budget that suited us. We derive both failures as properties of the estimators, and release the protocol, per-seed data, and a checklist whose first row is the one we needed: check that a headline effect is not produced by the instrument that measures it.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.