acceptodds
Under review as a conference paper at ICLR 2027

Scaling Dexterous Manipulation via Action Representation Pretraining

Abstract

Large-scale vision, language, and video pretraining has substantially improved the perception and generalization capabilities of robotic manipulation models, but action generation for complex dexterous manipulation remains constrained by the scarcity of robot demonstrations. Limited robot demonstrations must be used to learn both action representations and task-conditioned action distributions, making it difficult to train action representations sufficiently and, in turn, limiting the ability of downstream generative models to generate complex dexterous actions. To address this limitation, we propose ARPG (Action Representation Pretraining and Generation), which explicitly separates action representation learning from action generation: we first pretrain an action representation encoder on large-scale human manipulation data. We then freeze the encoder and train a compact generative backbone (130M parameters) to learn task-conditioned action distributions in the encoder's representation space. Experimental results show that ARPG achieves state-of-the-art performance on the DexJoCo benchmark and improves the average success rate on bimanual dexterous manipulation from 10.1% for π₀.₅ to 44.7%. In a nested data-scaling experiment, downstream performance improves as human-manipulation pretraining data increase from 10 to 500 hours. These results indicate that action representation pretraining provides an independent scaling dimension for improving complex dexterous manipulation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.