acceptodds
Under review as a conference paper at ICLR 2027

Vorch-Human: Unified Task Representation for Human-Centric Audio-Visual Generation and Long-Horizon Continuation

Abstract

Human-centric audio–visual generation remains fragmented across task-specific conditioning interfaces, limiting supervision sharing and consistent control across generation modes. To address this fragmentation, we present Vorch-Human, a unified framework built on a dual-stream audio–video diffusion transformer. Its unified task representation organizes target and condition tokens using role embeddings, temporal position types, and clean-token masks, bringing text- and image-conditioned generation, audio-driven human animation, and appearance–voice referenced joint generation into a common framework. Building on this representation, cross-task condition composition combines persistent appearance and voice references with rolling temporal context to extend both audio-driven and reference-conditioned generation to long-video continuation. We train the model in two stages: multi-task mixed training establishes shared generation capabilities, while subsequent context adaptation teaches the model to continue generation from historical audio and video. At inference, the temporal context is held fixed within each generation window and updated between windows, while persistent identity references maintain identity consistency throughout the sequence. Extensive experiments and human evaluations demonstrate competitive visual quality, identity preservation, voice similarity, and audio–visual synchronization, together with stable five-minute long-horizon video generation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.