CAD-WAM: Dexterous Manipulation Generation via Contact-aware World Action Model
Abstract
Generating 4D hand interaction trajectories from a single egocentric image and a language instruction requires anticipating how hand–object contact evolves across different hand regions. Most existing methods map vision and language directly to actions without explicitly imagining how the interaction unfolds, and their global conditions leave fine-grained contact under-constrained, so contact emerges only as an implicit by-product of global motion. We present CAD-WAM, a contact-aware world action model for dexterous generation which jointly imagines future egocentric video and hand actions, grounding trajectory generation in imagined interaction dynamics rather than static observations. Our key idea is to recast regional contact evolution from an implicit consequence of global motion into an explicit, predictable discrete prior for fine-grained refinement. Specifically, a coarse-to-fine dual-stream Mixture-of-Transformers (MoT) based network is proposed for both dexterous and imagination generation, where high-noise steps establish global interaction dynamics and a contact-aware action expert injects predicted regional contact skill codes at low-noise steps to refine finger articulation and contact transitions. The codes are learned by a spatiotemporal VQ-VAE over six palmar regions and predicted from the instruction and the initial contact state, requiring no future observation at inference. The anatomical contact atlas engine further extracts vertex-level contacts from egocentric videos into six anatomically defined regional atlases. The detailed experiments are conducted to show the effectiveness of our method. On HoloAssist and ARCTIC, CAD-WAM reduces MPJPE by over 39% and wrist translation error by over 30% relative to the strongest baseline. Code will be released.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.