acceptodds
Under review as a conference paper at ICLR 2027

WorldRIDE: Efficient Reward-Guided Distillation for Autoregressive World Models

Abstract

Autoregressive world models must combine low-latency generation with visual fidelity and reliable action control. Few-step distillation reduces sampling cost, but treating reward alignment as a subsequent post-training stage introduces another optimization loop. We present WorldRIDE, an efficient reward-guided distillation framework that improves a four-step autoregressive student during distillation itself, without additional post-distillation training. WorldRIDE combines three designs. Ordered reward-guided distillation gates action preferences by group-level visual quality and permits geometric refinement only among action-safe, action-matched candidates. Multi-anchor rollout reuses a single autoregressive trajectory to construct local candidate groups at multiple temporal anchors, amortizing prefix generation across training targets. Locally self-normalized distillation converts within-group rewards into detached, unit-mean weights on the distribution-matching update, retaining unweighted distillation when the quality gate fails. On WBench, WorldRIDE achieves an Overall score of 77.1. On VBench, our four-step model outperforms a 60-step causal model on all three aggregate scores, improving Total, I2V, and Quality by 0.526, 0.767, and 0.286 points, respectively. These results show that reward-guided distillation can improve world-model quality while reducing denoising steps, without a separate stage to recover performance after distillation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.