acceptodds
Under review as a conference paper at ICLR 2027

One Length Does Not Fit All: Unifying Video RL across Length Scales

Abstract

Following its success in aligning large language models with human preferences, reinforcement learning (RL) post-training has become a dominant paradigm for aligning video generation models with user preferences. However, existing approaches typically perform RL on a single, fixed video length: restricted to a single length scale, the learned improvements amount to local adaptations that fail to generalize across durations. This raises a fundamental question unexplored: how do the optimization dynamics of RL inherently differ across video length scales? In this work, we systematically reveal a severe Length-Scale Inconsistency in video RL. Since short and long videos place different emphasis on information dimensions during optimization, training under the same reward function over-exploits whichever dimensions dominate the reward at that length: static color and texture cues for short videos, whereas temporal shortcuts cause incoherent motion and abrupt transitions for long videos. To address this, we propose the Length-Aware Training Framework (LA-Framework), which unifies RL for videos of varying lengths through two core designs that directly resolve the underlying causes: (1) a length-conditioned timestep credit assignment mechanism that efficiently estimates reward sensitivity via sparse branching and shared prefixes, constructing a weight tree to dynamically allocate gradient budgets to the timesteps that actually determine the reward at each length; and (2) a multi-length joint training paradigm that normalizes advantages within identical length groups, allowing length-specific hacking directions to counteract each other while jointly reinforcing genuinely beneficial improvements. Extensive experiments across nine evaluation dimensions demonstrate that our framework effectively mitigates length-specific reward hacking and achieves high generation quality consistently across various video lengths.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.