One Token per Speaker: An Omni-Native, Exchangeable Entity Interface for Conversational Memory Routing
Abstract
Speaker-specific memory requires linking utterances to the correct person. We investigate whether an omni model can derive and use speaker identity from its own audio features without a separate speaker encoder. A learned pooler over frozen Qwen2.5-Omni features produces one continuous token per enrolled speaker. With LoRA adaptation, speaker identification for session-memory retrieval reaches 97.3%, compared with 24.8% for mean pooling of the same features. Without target-domain adaptation, accuracy reaches 74.3% on VoxCeleb1 and 58.3% on AliMeeting. Deterministic ordering preserves predictions when unchanged enrollment entries are reordered. Reliable identification does not, however, ensure effective use in dialogue. Under controlled candidate scoring, a continuous identity cue changes only one of 160 decisions relative to no identity. Providing the same resolved identity as text reduces episodes with wrong-person writes from 77.5% to 12.5%. The tested speaker-side and model-side adaptations do not close this gap. Over 30-turn meeting sequences, cosine matching over the same tokens roughly halves wrong-person writes relative to in-model use, although its errors persist longer. Mapping cosine-matching results to existing candidate embeddings matches the gold-identity reference without additional parameters or training. This result is limited to the tested single-token candidate setting; the free-generation check ran but did not establish safety, reaching 25 of 40 valid decisions against a pre-registered floor of 32. Together, these findings demonstrate useful omni-native speaker representations while identifying their connection to downstream decisions as a remaining challenge.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.