Learning to Act through Sparse 4D Queries
Abstract
Successful manipulation requires a policy to understand where things are and how they will move, yet action demonstrations leave this spatiotemporal structure implicit. Existing methods supply it through dense supervision: world action models predict every pixel of future frames, and 3D-aware vision-language-action (VLA) models align every token with teacher features, both at substantial training cost. We ask whether sparse, explicit supervision suffices. We introduce SparQ4D, which casts depth, 2D/3D point motion, multi-view correspondence, and hand-eye relations as a single D4RT-style question (where is a selected point, at a specified time, in a specified view?), which is answered by a lightweight query decoder attached to a pretrained VLA. Each training step samples only a small query budget, and the decoder is discarded after training, leaving the VLA's architecture and inference cost unchanged. Across two VLA backbones in simulation and the real world, SparQ4D consistently improves manipulation: it raises the success rate of from 77.6% to 92.6% on RoboTwin 2.0, reducing over 60% failures, and outperforms X-WAM with 20% of its training GPU-hours and faster inference. Probing further shows that the trained VLA itself learns recoverable 4D structure.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.