acceptodds
Under review as a conference paper at ICLR 2027

ExDiff: Diffusion as an Expressive Prior for Hand Gestures and Facial Expressions

Abstract

Expressive hand gestures and facial expressions provide essential non-verbal cues for understanding human intent in image-based human-computer interac- tion. However, existing pose priors often focus on full-body motion or model the hands and face as independent components, limiting their ability to capture the dependencies between hand articulation and facial expression. We propose EXD- IFF, an expressive diffusion framework for jointly modeling bilateral hand ges- tures, jaw motion, and facial expressions within a unified expressive latent space. The proposed framework introduces part-balanced diffusion with part-aware noise schedules, cross-part residual fusion, and geometry-aware denoising supervision to model fine-grained hand articulation, facial dynamics, and their mutual depen- dencies. To support heterogeneous training, EXDIFF learns part-specific hand and face diffusion priors, then trains a zero-initialized cross-part residual fusion head on joint hand–face data, using part-availability masks and confidence-aware con- ditioning. The learned prior supports expressive generation, missing-part comple- tion, and monocular hand–face mesh recovery. Experiments show that EXDIFF consistently improves realism, coverage, and diversity on ARCTIC, FreiHAND, NOW, and BEAT2. The code will be made available to the public.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.