acceptodds
Under review as a conference paper at ICLR 2027

SocialDiT: Social-Aware Diffusion Transformer for Human-Human Interaction Generation

Abstract

Two people performing the same action move differently depending on who they are and how well they know each other, yet text-driven dyadic motion generators condition on the action description alone. We study text-driven dyadic interaction generation additionally conditioned on per-actor Big Five personality and pair-level familiarity. We introduce SocialDiT, a flow-matching diffusion transformer that feeds the caption through text cross-attention and the social attributes through a Social-Aware Cross-person Attention (SACA) module, which modulates the attention between the two participants. We show that on Inter-X, where every subject pair carries a single fixed label, probes that predict these attributes from motion are confounded by pair identity. We therefore test whether the social input is used and controllable with interventions, a controlled familiarity probe, evaluator-free interaction geometry, and a user study. On a subject-disjoint split, SocialDiT reduces FID from to relative to the strongest socially conditioned baseline. It responds to interventions on its social input and to controlled changes in familiarity, reproduces interpersonal distances more closely than the baselines, and is preferred in a user study, including on matching the target familiarity. A controlled comparison of injection sites further shows that routing social attributes through cross-person attention best reproduces interpersonal geometry, suggesting that social context is most naturally injected where the two participants interact.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.