acceptodds
Under review as a conference paper at ICLR 2027

Sound Follows Contact: Streaming Audio-Visual Generation with Contact-Aware State Grounding

Abstract

Streaming audio-visual generation requires sound events to align with their corresponding visual causes. In rigid-body scenes, sound timing is strictly governed by physical contact, making it highly sensitive to subtle motion errors. Under autoregressive chunk-wise rollout, these errors accumulate across transitions and manifest as delayed, missing, or spurious audio events. We formalize this failure mode as contact-phase drift, which compounds over long horizons where contact timing degrades substantially faster than perceptual quality. To address this challenge, we present Tora-live, a state-grounded framework for streaming audio-visual generation in contact-dominated rigid-body scenes. Tora-live augments a few-step chunk-causal generator with a compact contact-aware hybrid state that combines dominant-object kinematics with explicit contact-phase variables. This state is maintained online through lightweight state propagation, contact-phase construction, and closed-loop correction from generated video, conditioning both modalities simultaneously through current-state modulation and short-term state-history attention. We further introduce a contact-focused rollout optimization objective that encourages generated audio onsets to align with scheduled contacts while suppressing missed and spurious responses. To evaluate this setting, we build ContactBench, a benchmark for contact-dominated audio-visual generation with contact-centric alignment metrics. Experiments show that Tora-live improves contact timing, audio-visual alignment, and long-horizon robustness over strong streaming baselines while maintaining competitive perceptual quality, with gains that increase consistently with rollout length.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.