VoiceMem: Streaming Dual-Brain Memory for Real-Time Interaction
Abstract
Real-time conversational systems require relevant memory within limited context and response-time budgets. We introduce VoiceMem, a streaming memory layer with factual and persona-affective retrieval branches, shared evidence identifiers, and optional entity attribution. An evolving schema–entity index adapts to recurring queries, while per-turn maintenance separately records immediate reactions and model-proposed persona claims. A streaming controller runs the same retrieval code before speech endpoint confirmation, overlapping preparation with the utterance so that only unfinished final ranking remains on the memory-critical path when preparation is complete. VoiceMem supports interchangeable backends and includes a teacher-supervised adaptation workflow, ChatMem-400K, and the audio-grounded ChatMem-Bench. Evaluations cover factual, persona, and audio-memory quality, context efficiency, component contributions, and three backend integrations. VoiceMem improves on reported factual and persona baselines, including a 21.43-point gain over Full-Context on LoCoMo, and exceeds MemOS by 14.20 percentage points on ChatMem-Bench. At the default LoCoMo operating point (K=5), the same text-evaluation setting achieves 90.4 with 430 memory tokens and 134 ms mean memory-system latency without streaming overlap. Streaming further hides preparation behind speech, connecting evidence organization with response-time memory preparation while separating memory management from the response model.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.