OmniMemory: Towards Unified Omnimodal Memory for Long-Horizon Interactive Assistants
Abstract
Long-horizon interactive assistants must repeatedly reason over audiovisual experiences that accumulate beyond finite context windows. Existing memory methods often prioritize visual observations and speech transcripts, underrepresenting richer auditory cues. Fragmented context, persistent interpretation errors, and perceptual information loss further limit memory reliability. We introduce OmniMemory, a training-free framework that builds evolving audiovisual memory with a unified omnimodal model. Progressive construction propagates scene and entity context across clips to build Episodic Memory, explicitly preserving visual events, speaker associations, utterance-aligned paralinguistic attributes, environmental sounds, and music. Retrospective auditing uses later evidence to revise earlier identity assignments and update affected records, while consolidation distills reusable knowledge into Core Memory. Query-aware routing reuses textual memories and selectively revisits source audiovisual clips in Native Memory when additional perceptual evidence is needed. Evaluations on M3-Bench-Robot, M3-Bench-Web, and VideoOdyssey-AV demonstrate the highest overall accuracy among evaluated systems, with substantial gains over memory baselines on Sound and Music questions. Ablations demonstrate the benefits of auditory fields, retrospective auditing, and complementary memory sources, while efficiency analysis shows a favorable trade-off among accuracy, construction overhead, and query costs.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.