Sensorimotor Alignment: A Training-Free Predictor of Robot Policy Success
Abstract
Robotic foundation models are typically evaluated through costly closed-loop rollouts, which motivates cheap offline proxies for model capability. The Platonic Representation Hypothesis (PRH) suggests that as representations improve, different modalities such as vision and language converge toward a shared statistical model of reality. We ask whether motion trajectories form another such modality: as robot models become more capable, should their sensory representations converge toward the structure of physical action? To examine this, we propose **Sensorimotor Alignment** (), a training-free offline metric that measures how well a model’s sensory representation aligns with the geometry of physical action by comparing sensory and trajectory similarity structures. Crucially, consistently and strongly correlates with downstream robot task success across embodied model families, making it a low-cost evaluation proxy. It can also identify promising VLM backbones even before costly VLA fine-tuning. Moreover, we show that , as a diagnostic lens, can explain how design choices shape where control-relevant structure is represented in the models. More broadly, offers to bring PRH into embodiment, turning sensorimotor convergence into a measurable object of study.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.