acceptodds
Under review as a conference paper at ICLR 2027

FLARE: Improving Long-Horizon Fidelity in Autoregressive Video Generation via Factorized Latent Reward Feedback

Abstract

Autoregressive video diffusion with streaming training enables real-time long-form generation, but error accumulation across chunks makes long-horizon consistency a fundamental challenge. Existing streaming distillation objectives keep generated chunks on the natural-video manifold, but provide limited guidance for resolving quality and alignment ambiguities during autoregressive rollout. As a result, long rollouts often suffer from semantic drift, visual quality decay, and motion degradation. We introduce FLARE, which closes this gap through a latent preference reward that scores video latents directly via a DiT backbone, equipped with three learnable reward queries that attend to backbone features and route to independent heads, allowing a single video pair to express conflicting preferences across aspects. The reward is trained in two stages: a generic Bradley-Terry preference SFT, followed by a consensus-guided latent preference adaptation stage in which multiple pixel-space evaluators serve as weak head-wise annotators and only cross-evaluator-agreed comparisons are retained. The trained reward is then integrated as an online, differentiable per-chunk signal alongside DMD distillation, providing aspect-specific preference guidance during streaming optimization while retaining distribution-level supervision from DMD. Experiments on long-video generation show that FLARE improves generation fidelity over strong autoregressive baselines, and ablations verify the effectiveness of the latent reward design.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.