PRIVILEGED FORESIGHT DISTILLATION: AMORTIZED FUTURE CORRECTION FOR WORLD ACTION MODELS
Abstract
World action models jointly predict future video and action during training, raising an open question about what role the future-prediction branch actually plays. A recent finding shows that this branch can be removed at inference with little to no loss on common manipulation benchmarks, suggesting that future information may act merely as a regularizer on the shared visual backbone. We propose instead that joint training induces an action-conditioned correction that privileged future observations impose on action denoising, and that current-only policies capture this correction only partially. Making the account precise, we formulate privileged foresight as a residual in the action-denoising direction—the difference between what a model predicts given the true future and what it predicts given only the current frame—and introduce Privileged Foresight Distillation (PFD), which transfers this residual from a training-time teacher into a small adapter on a current-only student. The teacher and student share the same backbone and differ only in the attention mask over video tokens; future video is never generated at inference. Controlled experiments support that this gain reflects a future-conditioned correction rather than added capacity: with the total teacher weight held fixed, routing it directly rather than as a residual drops the policy below the undistilled baseline. Empirically, PFD improves over Fast-WAM on LIBERO and RoboTwin manipulation benchmarks while preserving the current- only inference interface with only a slight adapter-induced latency overhead. This view reframes the role of future information in world action models: not as a target to predict, nor as a branch to discard, but as a compressible correction to be distilled—one that the privileged path carries even when it is the worse predictor.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.