acceptodds
Under review as a conference paper at ICLR 2027

Ex-Omni-2D: Expressive Omni-Modal Dialogue Models with Native Visual Presence

Abstract

Omni-modal dialogue models can understand multimodal inputs and synthesize spoken replies, but a spoken answer still leaves the agent visually absent. We introduce Ex-Omni-2D, a framework that answers a multimodal query with coordinated text, personalized speech, and reference-conditioned video. The LLM first generates a structured Visual Thought Plan (VTP) followed by the response text. The speech generator then produces multi-codebook speech units from response-side representations. These speech units are decoded into audio and aligned with video frames, giving the speech and avatar modules a common timing signal while allowing them to learn from different data sources. The full-sequence video generator realizes the framework by conditioning on reference appearance, VTP semantics, and frame-aligned speech units. This is the configuration used for our main quality evaluation. For latency-sensitive settings, we additionally adapt it for few-step, block-causal streaming generation with a dedicated reference-image sink and speech-aligned chunks. This provides incremental output with lower startup latency than full-sequence decoding.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.