acceptodds
Under review as a conference paper at ICLR 2027

OmniMemory: Towards Unified Omnimodal Memory for Long-Horizon Interactive Assistants

Abstract

Long-horizon interactive assistants must repeatedly reason over audiovisual experiences that accumulate beyond finite context windows. Existing memory methods often prioritize visual observations and speech transcripts, underrepresenting richer auditory cues. Fragmented context, persistent interpretation errors, and perceptual information loss further limit memory reliability. We introduce OmniMemory, a training-free framework that builds evolving audiovisual memory with a unified omnimodal model. Progressive construction propagates scene and entity context across clips to build Episodic Memory, explicitly preserving visual events, speaker associations, utterance-aligned paralinguistic attributes, environmental sounds, and music. Retrospective auditing uses later evidence to revise earlier identity assignments and update affected records, while consolidation distills reusable knowledge into Core Memory. Query-aware routing reuses textual memories and selectively revisits source audiovisual clips in Native Memory when additional perceptual evidence is needed. Evaluations on M3-Bench-Robot, M3-Bench-Web, and VideoOdyssey-AV demonstrate the highest overall accuracy among evaluated systems, with substantial gains over memory baselines on Sound and Music questions. Ablations demonstrate the benefits of auditory fields, retrospective auditing, and complementary memory sources, while efficiency analysis shows a favorable trade-off among accuracy, construction overhead, and query costs.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.