ReactMotion: Generating Reactive Listener Motions from Speaker Utterance
Abstract
When a friend says I finally got the job, we lean in, clap, or raise a hand; when the news is bad, we respond differently. A listener's body answers what is said, and we study how to generate this answer from the speaker's audio, transcript, and emotion. Two obstacles stand in the way: a recording shows only one of the many reactions an utterance can elicit, and a generic nod looks acceptable next to almost any utterance, so evaluation must measure whether a motion depends on its input. We address both. ReactMotionNet builds one-to-many supervision at scale: it curates an inventory of listener reactions, writes utterances that would elicit each one, cross-pairs conditions with candidate motions, and grades every pair with three language-model scorers calibrated on human judgments, yielding 85,142 condition–motion pairs, 71,286 of them supportive or socially appropriate. ReactMotion is a multimodal language model that predicts body-and-hand motion tokens directly from the speaker condition, without an intermediate motion caption, trained on the supportive pairs and refined with Direct Preference Optimization. Gain scores how much better a motion fits its own utterance than unrelated ones, so an input-blind generator scores zero. On 358 held-out utterances, ReactMotion reaches 0.3392 Gain after supervised fine-tuning and 0.3790 after preference optimization, against at most 0.1157 for twelve caption-mediated pipelines, and human raters prefer it to retrieval and cascade baselines.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.