Sonic-AMG: Streaming Human Motion Generation from Spatial Audio
Abstract
Modeling human reactions to spatial audio requires understanding both what is heard and where it comes from. Prior work generates motion offline from complete acoustic and spatial input sequences. Humans, however, react to sounds without knowing what they will hear next. We present Sonic-AMG, the first autoregressive framework for generating human motion from streaming spatial audio. The model encodes acoustic content and body-relative source geometry as separate tokens. A causal transformer with a flow head generates continuous motion latents, which a causal decoder reconstructs into motion segments. Teacher-forced training uses motion histories and body-relative source locations derived from recorded motion, whereas streaming inference relies on generated motion for both. We introduce *rehearsal training*, which combines full-sequence teacher-forced supervision with supervised prediction on histories constructed through short autoregressive self-rollouts, improving robustness to accumulated prediction errors. To keep spatial conditioning consistent with generated motion, we propose *spatial feedback*, which recomputes body-relative source locations from generated root poses during both rehearsal training and inference. On SAM, Sonic-AMG achieves state-of-the-art FID and R-precision. It further supports real-time online generation of coherent reactions across sound events and source locations. Our code and model will be released upon acceptance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.