acceptodds
Under review as a conference paper at ICLR 2027

OmniPersona: Persona-Embodied Agents for Multimodal Conversation Interaction

Abstract

Current interactive talking avatar models remain fundamentally unaware of dialogue context, requiring explicit textual or audio driving signals for each response, thereby functioning more like video renderers than autonomous avatar agents. Inspired by recent advances in unified understanding and generation frameworks, we present OmniPersona, an agentic model that integrates a pretrained multimodal large language model (MLLM) with an adapted joint audio-video generator. Our framework unifies dialogue understanding with synchronized audio-video avatar generation within a single architecture. Central to this design are learnable queries that compress historical multimodal context into a fixed-size non-verbal memory. Together with verbal content, this memory enables our generator to produce more responsive avatar behaviors, particularly in the listening state, which remains underexplored in prior work centered on talking scenarios. Additionally, we incorporate a suite of anti-drifting techniques, including a masked fusion attention design, to support persona consistency (audio-visual identity and synchronization) across multi-turn conversations. This attention design enables customized voice conditioning without compromising audio-video consistency. Extensive experiments on three datasets and human evaluation demonstrate that OmniPersona achieves superior performance.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.