acceptodds
Under review as a conference paper at ICLR 2027

Anamora: Reference-Grounded Streaming Audio-Video Generation with Entity-State Memory

Abstract

Streaming joint audio-video generation is an emerging paradigm for continuous audiovisual synthesis, opening new possibilities for immersive, real-time interaction. Stable subject identities and coherent interactions are essential to this process, yet current streaming approaches struggle to maintain consistent subject and scene states when shots change or multiple subjects interact. Stable subject integration and sustained multi-subject interaction are therefore central to coherent, personalized streaming experiences. We present Anamora, a unified framework for reference-grounded multi-shot streaming audio-video generation with entity-state memory. Anamora keeps shot-specific visual references accessible to maintain subject identities while allowing reference updates during generation. A state-aware entity memory selects history compatible with the current requirements and shares a fixed capacity across subjects, supporting state continuity across scene transitions and reappearances. Extensive experiments demonstrate that Anamora achieves high-fidelity subject injection and cross-shot consistency, highlighting the importance of reference grounding and entity-state memory in streaming interactive audio-video generation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.