acceptodds
Under review as a conference paper at ICLR 2027

DeForcing: Efficient Rolling-Free Training and Distillation Framework for Streaming Real-time Avatar Generation

Abstract

Streaming real-time avatar generation is important for live video conferencing, virtual agents, and interactive digital humans, where visual responses must be synthesized causally and with low latency while preserving identity, motion co- herence, and audio–visual synchronization. We present **DeForcing**, a rolling-free training and distillation framework that converts a bidirectional diffusion avatar model into a streaming causal generator. For training, the framework uses op- tional causal warm-up (Causal SFT or CFG distillation) followed by DMD, with- out pre-generated ODE trajectories from expensive multi-step teacher inference or sequential self-forcing rollouts. Under noisy-cache inference, noise-aligned his- tory and a fixed block-causal mask make denoising invariant to temporal chunk- ing. The same model therefore supports autoregressive (AR) rollout for low- latency streaming, multi-token prediction (MTP) for parallel denoising of mul- tiple temporal chunks, and single-pass prefill for higher throughput, without re- training. We further combine hybrid local–global temporal attention across layers with attention-sink caching under a fixed KV-cache budget to reduce attention cost and enable long-horizon streaming. Instantiated in our audio-driven avatar system with reference-conditioned generation, the framework empirically supports the equivalence of these decoding schedules and demonstrates stable appearance and scene layout over long-horizon continuation. The resulting model maintains a fa- vorable balance among visual quality, audio–visual synchronization, and identity preservation relative to the bidirectional teacher.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.