acceptodds
Under review as a conference paper at ICLR 2027

Efficient Post-Training for 3D-Consistent Video Generation via Optimal Transport Reward Balancing

Abstract

Reinforcement learning (RL) can improve the 3D awareness of video diffusion models without architectural modifications. However, existing methods face three challenges: reconstruction rewards favor limited camera motion, heterogeneous reward distributions cause imbalanced optimization, and policy updates over many denoising steps are computationally expensive. We propose 3D-Balanced-R1 to address these issues. First, a visibility score measures inter-frame overlap and reweights the reconstruction reward according to reconstruction difficulty. Sec- ond, optimal transport aligns different rewards to a common Gaussian distribu- tion, reducing domination by individual rewards or outliers. Third, we update the policy at only two denoising steps, using an early step for camera control and a late step for video quality and 3D consistency. On Wan2.1-1.3B, 3D-Balanced- R1 achieves a 3.53× training speedup over World-R1. With the same number of training epochs and comparable video quality, 3D-Balanced-R1 improves the pre- trained model’s 3D score by 9.6%, compared with 6.3% for World-R1. Under a similar training-time budget, the gain increases to 14.3%, or 2.3× that of World- R1. Similar improvements are also observed on the larger Wan2.2-5B model.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.