acceptodds
Under review as a conference paper at ICLR 2027

Disentangled Representations for Zero-shot Behavioral Style Transfer

Abstract

Human behavior is exhibited through bodily motion that not only achieves its intent-driven or communicative goals, but also demonstrates stylistic and idiosyncratic patterns. Existing approaches for generative motion synthesis learn to transfer these styles to novel motions by learning adapters or designing task-specific guidance objectives. However, disentangling style and content within the core motion representation remains an open problem. We present DisCodec, a general framework for building disentangled representations of style and content in human motion, which can be used for zero-shot style transfer at the decoding stage. Our approach learns these representations through a neural motion codec trained with a Variational Information Bottleneck and SNR-adaptive Transform Quantization to disentangle style from the motion's action content. Two streams of encoders decompose any reference motion into style and content latents, which are then combined to reconstruct the motion, while also ensuring that these latents represent mutually distinct information about the motion. Therefore, any generative motion synthesis task learned using the proposed representation not only inherits zero-shot style transfer capability but can also sample styles from the learned style subspace. We benchmark the style transfer performance against existing approaches and independently demonstrate applications of DisCodec representation in text-to-motion synthesis and co-speech gesture generation tasks. Reader is urged to watch accompanying video demo.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.