Injected but Not Accumulated: Measuring and Stabilizing Synchronization in Training-Free Streaming Audio-Video Generation
Abstract
Joint audio-video diffusion models generate video with synchronized sound, but only within native windows of a few seconds. We study their training-free streaming extension: a causal rolling window whose memory and per-second cost do not grow with duration. Concurrent work trains streaming generators on the premise that extension makes audio-visual synchronization drift, and a concurrent benchmark reports endpoint drift of mixed sign across trained systems without resolving its mechanism. We present a measurement protocol that tracks synchronization over generation time with a calibrated offset estimator whose sensitivity is audited on every compared system, sample-accurate seam diagnostics, content-validity gates, and a stratified prompt suite with a negative control. The measurement overturns the premise. Naive streaming injects a conditioning-anchor error of −120 ms at 91% of window seams, as a closed-form analysis of the conditioning arithmetic predicts to within 2.5 ms across ten operating points, yet the errors do not accumulate: 4.3 s of injected error over two minutes leaves a median realized offset of 0.23 s, and offset curves stay flat for five minutes, because each window's joint denoising re-synchronizes audio to video. The same protocol yields three training-free stabilizers, modality-aligned conditioning arithmetic, audio-lead conditioning, and cross-window fixed noise, whose combination removes the seam fingerprint and improves synchronization over both naive streaming and an offline joint-denoising baseline on a 200-prompt benchmark (0.253 vs. 0.285 and 0.437 s expected offset), at matched instrument sensitivity and no measured visual-quality cost. On a second backbone with a different fusion design and conditioning path, naive continuation does degrade synchronization after the first window, and the fixed-noise stabilizer recovers part of the loss, with prompt-dependent magnitude. Cross-window subject identity remains open: our system re-rolls identity at 72% of seams, where the offline baseline, at full-duration latency, almost never does. We release the protocol, the prompt suite, and code.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.