acceptodds
Under review as a conference paper at ICLR 2027

Dual-Support Online Policy Optimization for Few-Step Autoregressive Video Models

Abstract

Few-step causal autoregressive (AR) video models enable efficient streaming generation through timestep distillation and KV-cached causal inference, making online optimization directly on the final deployed model increasingly practical. However, post-distillation optimization introduces a deployment-support mismatch: the training updates may be estimated at timesteps and causal contexts that differ from those actually encountered by the few-step streaming policy at inference. We identify two complementary sources of this mismatch. The distilled timestep support is the discrete timestep set retained by the few-step sampler, while the streaming context support is the distribution of causal contexts induced by the model's own autoregressive rollout. Based on this observation, we propose a dual-support online policy optimization framework that performs relative forward-process updates only on distilled timesteps and optimizes candidate continuations under policy-generated long-horizon contexts. Historical KV states are detached so that optimization remains local while context exposure remains long-horizon. Across multiple distilled AR video backbones and generation horizons, our method consistently improves preference-oriented quality and motion dynamics while preserving the original few-step inference procedure. Our results show that effective post-distillation alignment depends not only on the optimization objective, but also on matching the state distribution encountered by the final generative policy at inference.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.