acceptodds
Under review as a conference paper at ICLR 2027

Resilient Streaming Video Editing

Abstract

Existing streaming video editors can exhibit ruminative behavior, over-relying on historical frames even as errors accumulate in the generation history. To address this limitation, we present JoyAI-Video-Edit, a resilient 16B-parameter autoregressive diffusion model for source-faithful, low-latency streaming editing. Specifically, we introduce Resilient-DMD, which selectively strengthens teacher-side source guidance during distillation when relative source attention falls below matched baselines, yielding a two-step generator. Long-Horizon Autoregressive Distillation (LHAD) extends this adaptive supervision to deeper rollout states while bounding activation memory. Our streaming engine combines bounded transformer state, single-frame VAE state handoff, and asynchronous execution to reduce end-to-end overhead. To support training and sustained-stream evaluation, we construct an editing data pipeline and introduce LongV2VBench, built on one-minute videos. Automatic and human evaluations show that JoyAI-Video-Edit outperforms the evaluated streaming baselines and remains competitive with strong offline editors. The system delivers end-to-end editing at 34.87 FPS on a single B200 GPU, with support for RTX 5090 and RTX PRO 6000.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.