acceptodds
Under review as a conference paper at ICLR 2027

TempoTrack: Temporally Grounded Multi-Track Control for Streaming Joint Audio-Video Generation

Abstract

Interactive joint audio-video generation requires users to schedule independently timed events and introduce new instructions as generation proceeds. Existing global conditioning and segment-level prompt switching offer limited support for dynamically arriving controls with asynchronous, overlapping intervals. We introduce TempoTrack, a framework for online, temporally grounded multi-track control in streaming joint audio-video generation. Its core mechanism, Hierarchical Context Routing (HCR), connects event-level instructions to blockwise generation through coarse-to-fine conditioning. An online prompt queue admits newly received instructions at block boundaries and selects events relevant to each causal window, while Temporal Attention Bias (TAB) aligns instruction intervals with audio-video query positions on a shared physical timeline through timestamp-dependent cross-attention biases. To enable few-step streaming generation, we directly compose the temporally controllable causal model with a pretrained distillation LoRA and apply self-forcing distribution matching under self-generated histories, without an additional ODE-based initialization stage. We further introduce MTBench to evaluate multi-track instruction adherence, temporal alignment, and joint audiovisual quality. Experiments demonstrate asynchronous, overlapping, and dynamically arriving controls alongside strong identity consistency and audiovisual synchronization, enabling streaming audiovisual content to be incrementally directed through independently scheduled events.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.