ISAudio: Interactive Spatial Audio Generation for World Models
Abstract
Spatial audio is essential for world models to simulate immersive interactions with action-responsive auditory feedback. Unlike video-to-audio synthesis from fully observed clips, this interaction loop requires streaming spatial audio that responds to listener actions and predicts upcoming sounds before the corresponding visual observations arrive. In this work, we introduce ISAudio, an action-conditioned framework for streaming binaural audio prediction. To address data scarcity, we construct spatial audiovisual data from panoramic recordings with controlled listener trajectories and first-person gameplay with aligned action signals. For efficient streaming, ISAudio combines autoregressive flow matching in a spatial latent space with causal consistency distillation, enabling few-step prediction conditioned only on observed visual prefixes, audio history, and current actions. To preserve semantic, temporal, and spatial consistency, we propose Contrastive Spatial Audio-Video Pretraining (CSAVP), which learns representations that capture changes in listener viewpoint through same-scene trajectory contrasts. We further refine the distilled generator with dual-granularity Group Relative Policy Optimization (GRPO), combining clip-level perceptual rewards with chunk-local event and spatial feedback. Experiments show that ISAudio achieves an S-CLAP score of 0.787, compared with 0.664 for the strongest baseline, and enables real-time streaming inference with a real-time factor of 0.53. Together, these results highlight interactive spatial audio generation as an important step toward richer multimodal world models.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.