SocialDiT: Social-Aware Diffusion Transformer for Human-Human Interaction Generation
Abstract
Two people performing the same action move differently depending on who they are and how well they know each other, yet text-driven dyadic motion generators condition on the action description alone. We study text-driven dyadic interaction generation additionally conditioned on per-actor Big Five personality and pair-level familiarity. We introduce SocialDiT, a flow-matching diffusion transformer that feeds the caption through text cross-attention and the social attributes through a Social-Aware Cross-person Attention (SACA) module, which modulates the attention between the two participants. We show that on Inter-X, where every subject pair carries a single fixed label, probes that predict these attributes from motion are confounded by pair identity. We therefore test whether the social input is used and controllable with interventions, a controlled familiarity probe, evaluator-free interaction geometry, and a user study. On a subject-disjoint split, SocialDiT reduces FID from to relative to the strongest socially conditioned baseline. It responds to interventions on its social input and to controlled changes in familiarity, reproduces interpersonal distances more closely than the baselines, and is preferred in a user study, including on matching the target familiarity. A controlled comparison of injection sites further shows that routing social attributes through cross-person attention best reproduces interpersonal geometry, suggesting that social context is most naturally injected where the two participants interact.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.