acceptodds
Under review as a conference paper at ICLR 2027

HIDDENFACE: STREAMING CONVERSATIONAL FACIAL MOTION FROM HIDDEN STATES OF A LARGE AUDIOLANGUAGE MODEL

Abstract

Full-duplex large audio-language models (LALMs) enable voice agents to take part in conversations, but embodying these agents as avatars also requires facial motion that stays aligned with their speech and continues through listening and turn transitions. However, audio-driven methods must reprocess the generated speech in a cascade, incurring processing and synchronization overhead. Methods that drive facial motion from LALM internal states mitigate this, but they train the LALM jointly, rely on additional generated speech units, or delay the output to refine motion with later hidden states. We propose HiddenFace, which generates conversational facial motion using only the hidden states already computed by a frozen LALM (Voila) during speech generation. HiddenFace combines Voila's audio and text hidden states with its own motion-token history to autoregressively predict the next facial motion token, finalizing each 160 ms block as soon as its hidden states are available, without waiting for future speech. Under matched 160 ms block-causal conditions, HiddenFace matches or exceeds the lip accuracy of audio-driven baselines on held-out real dyadic conversations and adds only a small, constant per-step cost to speech generation. At the dialogue level, it is closer to real human facial behavior than any baseline, even those using the full audio, and receives the highest mean ratings in a human study.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.