Single Stream, Single Step: Simplifying 3D Human–Human Interaction Generation
Abstract
Generating realistic 3D human–human interactions requires plausible individual motion and coherent coordination between two individuals. Existing approaches often use separate per-person streams or specialized interaction modules, while some additionally rely on learned motion autoencoders. These designs complicate the architecture or training pipeline, and iterative sampling adds inference cost. We introduce InterFlow, an end-to-end autoregressive flow-matching model that simplifies interaction generation using a single-stream diffusion transformer. InterFlow operates directly in motion space without a variational autoencoder, representing the motions of both individuals as a unified token sequence and modeling temporal and interpersonal dependencies through shared self-attention. Its fixed-window formulation enables blockwise generation with a computational cost per block independent of the total generated sequence length. We further introduce FastInterFlow, which combines rollout-based score distillation with forward-kinematics-based adversarial distillation to obtain a one-step generator from the multi-step teacher. FastInterFlow retains the autoregressive formulation and generates each motion block in a single network evaluation. Empirical results demonstrate the effectiveness of InterFlow’s streamlined architecture for modeling interactions. Together, the two models provide a simple framework combining unified motion modeling with efficient blockwise autoregressive generation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.