acceptodds
Under review as a conference paper at ICLR 2027

Aligning the Clocks: Alignment-Aware Causal Distillation for Streaming Audio-Driven Avatars

Abstract

Streaming audio-driven avatars require a video generator to produce each new chunk using only an image identity, past visual context, and a temporally bounded audio signal. Existing diffusion distillation recipes largely treat this problem as model compression. We argue that causal avatar distillation has an additional failure mode: the clocks of waveform time, audio features, video frames, VAE latents, and autoregressive blocks are not automatically aligned. This mismatch is amplified when a bidirectional teacher is distilled into a causal student, and can manifest as lip lag, weak mouth motion, chunk-boundary discontinuities, identity drift, and exposure drift. We introduce , an alignment-aware causal distillation framework combining causal projection of teacher trajectories, explicitly bounded audio visibility with monotonic audio–latent alignment, mouth-dynamics-aware score matching, and partition-invariant chunk consistency. The design preserves the original OmniAvatar audio-conditioning path rather than introducing an artificial audio cross-attention module. Our current implementation provides a reproducible baseline suite over OmniAvatar and CausVid. Preliminary results show that ODE initialization improves selected visual and synchronization dimensions but does not dominate direct causal DMD on every metric. We therefore separate completed evidence from the proposed controlled study, which tests whether temporal alignment, rather than ODE initialization alone, is the central bottleneck.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.