acceptodds
Under review as a conference paper at ICLR 2027

CineForcing: Streaming Multi-Shot Audiovisual Generation for Long Storytelling

Abstract

Humans experience the world as a continuous, synchronized stream of sound and vision. While modern generative models excel at synthesizing brief audiovisual clips, they struggle to sustain this continuity over long horizons. During extended generation or scene transitions, models suffer from identity drift: characters alter their appearance and voices lose their distinctive timbre. Tracking visual frames is insufficient, since acoustic identity cannot be recovered from pixels alone. We present **CineForcing**, a memory-conditioned streaming framework for long-horizon audiovisual generation. Instead of carrying an ever-growing buffer of raw tokens, CineForcing summarizes completed shots as compact committed memory states that pair a visual anchor with speech-filtered acoustic evidence from the same shot. Student generation is causal at two timescales: audiovisual blocks are generated sequentially within a shot, while each new shot reads only state committed by completed shots. Memory-conditioned supervised fine-tuning initializes the backbone; parallel audiovisual teacher forcing and two-pass distribution matching then train causal streaming with ground-truth and self-generated *within-shot* prefixes, respectively, under strict-past audiovisual memory. We evaluate CineForcing on 30-shot long-story audiovisual rollouts. On 100 stories comprising 3,000 shots, CineForcing achieves the highest Voice and CAVP scores and the lowest audiovisual desynchronization among the compared methods. It improves SpeechAcc from **0.4224 to 0.7583** over OmniForcing and, with four sampling steps, achieves a measured generation latency of **4.557 seconds per clip at 480p**. A matched inference-time ablation shows that audiovisual memory increases SpeechAcc from **0.0037** with both memory inputs zeroed to **0.7583**, while achieving the strongest late-shot visual consistency and synchronization among memory variants. Video demos: [cineforcing-demos.pages.dev](https://cineforcing-demos.pages.dev/).

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.