acceptodds
Under review as a conference paper at ICLR 2027

StreamPilot: Agentic Inference for Long-Horizon Streaming Video Generation

Abstract

Streaming video generation models produce videos autoregressively. This leads to two challenges that all streaming video generation models share. The first is the drift: errors in the generated content accumulate until the picture collapses. The second is long-horizon consistency: as the video grows and the prompt keeps changing, the model does not remember what it generated, and may forget relevant content or recall the wrong content. We propose StreamPilot, a training-free agentic inference framework in which a VLM agent that reads frames fills the long-term cache during inference. At each instruction change and at fixed intervals, the agent reads the current instruction and the latest frames, decides which earlier chunk comes back as a reference and checks each new chunk before it is stored. References are stored as latents and recomputed where they are loaded. Without any training, StreamPilot achieves better long-horizon consistency than the base models on three streaming video models and four public benchmarks, showing that an agent can manage the memory of a streaming video model at inference time. On Causal Forcing, it raises the share of MemFlow clips showing the requested scene from 29.0% to 73.5% and cuts the videos that collapse after a scene returns from 58% to 13%. It also lets Self Forcing surpass LongLive in memory without any training, although LongLive is tuned for long videos.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.