Beyond Future Prediction: Denoising as Generative Adaptation for Robot Control
Abstract
Pretrained generative Diffusion Transformers (DiTs) capture rich pixel-level visual and language-conditioned structure through large-scale image and video generation training. A growing line of robot policies builds on this generative prior through generative visual and action co-training, commonly instantiated as future visual prediction, yet recent methods increasingly move this prediction out of the inference and keep it only for co-training. This shift leaves open a more basic question, what a pretrained generative DiT actually contributes to action learning, and how this prior should be adapted for control. We introduce NowWAM, a future-target-free co-training formulation that denoises the current observation and predicts robot actions from the same visual stream, directly coupling the native generative objective to the action-facing representation across the denoising trajectory. Under matched controlled settings, past and future visual targets perform comparably, whereas restricting training to the clean endpoint substantially reduces robustness. This suggests that a separate future target is not essential for generative adaptation, but the continuum of denoising states remains an effective interface for control. On LIBERO-Plus, NowWAM reaches 87.7% with FLUX2-Klein, improving over the future-target co-training baseline by 6.1 points while halving training visual tokens (784392) and reducing step time from 2.85 s to 1.63 s, a 1.8 speedup. With the pure text-to-image Z-Image backbone, NowWAM still reaches 87.8%, confirming that strong control adaptation does not depend on video generation or image-editing backbones.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.