acceptodds
Under review as a conference paper at ICLR 2027

TurboT2VA: Fast Large-Scale Text-to-Video-Audio Generation via Score-Regularized Consistency Distillation

Abstract

Joint text-to-video-audio generation produces synchronized visual and acoustic content, but the long sampling trajectories and heterogeneous multimodal computation of large models make inference prohibitively expensive. We present TurboT2VA, a distillation and inference framework for accelerating a 19B-parameter joint video-audio model. Compared with discrete-time consistency, continuous-time consistency avoids finite-step discretization errors in the training target and provides more faithful supervision along the teacher trajectory, making it particularly attractive for few-step distillation. However, scaling continuous-time consistency to a 19B joint T2VA model is challenged by modality-imbalanced optimization, numerically fragile trajectory-tangent estimation, and the quality–diversity trade-off. TurboT2VA addresses these issues with per-modality normalization and a progressive dCMsCMsCM+DMD curriculum, where dCM establishes a stable initialization, sCM more faithfully learns the continuous teacher trajectory while preserving diversity, and DMD improves perceptual quality. On LTX-2, four-step distillation reduces generator latency from 50.52s to 2.51s at the standard evaluation resolution of 512768, achieving a 20.1 speedup while maintaining strong visual quality, audio fidelity, diversity, and video-audio synchronization. We further develop an architecture-aware inference stack that combines guarded W8A8 and fused operators, padded-text compaction, and modality-aware sparse attention while preserving dense cross-modal and text-conditioning paths without retraining. Under the high-resolution deployment setting at 10241792, the complete stack reduces generator latency from 318.74s to 5.83s on one NVIDIA H20, achieving a 54.67 generator-only speedup over the original teacher.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.