BEYOND RANKING: MAKING LONG-TERM MULTI- MODAL MEMORY ADDRESSABLE
Abstract
Long-term multimodal memory is useful only if a future query can find the evidence it needs. Yet storing the right observation does not make it easy to retrieve, especially when evidence is visual, surrounded by similar memories, or distributed across multiple memory units. We call this problem the addressability gap. Across three long-term multimodal memory benchmarks, we study this gap through three retrieval challenges: finding image-bearing evidence, distinguishing the required observation from similar memories, and recovering the complete evidence set. We introduce CACE, an addressability-oriented framework with two components. At construction time, Contrastive Addressability Encoding (CAE) describes each observation in the context of visually similar memories, making its distinguishing details searchable while preserving the original observation as evidence. At retrieval time, Competitive Evidence Assignment (CEA) preserves individual query tokens and competitively assigns each token's matching credit across candidate memories, allowing different parts of the query to recruit different textual or visual evidence under a fixed retrieval budget. Experiments on MemEye, SMMBench, and Mem-Gallery show substantial improvements in evidence retrieval across all three challenges, with corresponding gains in downstream answer quality across multiple reader models. These results suggest that long-term multimodal memory should be designed not only to store relevant experience, but to make that experience addressable by future queries.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.