acceptodds
Under review as a conference paper at ICLR 2027

WoRL: Efficient Reinforcement Learning Framework for Video World Models with Next-Video Reward Modeling

Abstract

While video world models have achieved remarkable progress, reinforcement learning (RL) remains largely underexplored for their post-training. RL offers a promising way to improve these models through self-generated rollouts without real and long videos with interactive signal, yet scaling it in autoregressive video generation remains challenging due to expensive long-video optimization, the lack of history-aware reward models, and limited exposure to accumulated consistency failures during standard rollouts. We introduce **WoRL**, an end-to-end RL framework that addresses these challenges through three components. **(1) Efficient Train–Test-Matched RL.** We eliminate the intra-clip train–test mismatch in autoregressive DiffusionNFT and introduce Parallel Clean-Prefix Attention, context-aware policy scheduling, and distributed execution to scale full-attention RL to large models and long video contexts. **(2) Next-Video Reward Model (NVRM).** NVRM evaluates each generated clip conditioned on its preceding video history, enabling history-aware preference optimization over long horizons. We further train the model with a graph-based data construction pipeline that builds 312.4k video-continuation preference pairs without human expert annotation. **(3) Revisit-Centered Training.** We construct trajectories that revisit observed scenes to expose accumulated inconsistencies and introduce rewind regularization to improve learning under limited generation diversity. WoRL achieves a **9.2× training speedup** over the naive implementation. NVRM reaches **74.4% consistency agreement** and **82.4% physical-plausibility agreement** with blinded human judgments, outperforming HPSv3 and 3D Score on both criteria. Across action-conditioned LingBot-World-Fast-14B and text-conditioned Helios-Distilled-14B, WoRL consistently improves long-term consistency, controllability, and physical plausibility, yielding a **37.5-percentage-point human-preference margin**. The qualitative results are best visualized as videos in the supplementary material.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.