LILIA: Benchmarking Long-Term Memory for Multi-User, Multi-Device Assistants
Abstract
Existing benchmarks have advanced long-term memory evaluation, yet provide limited coverage of interactions in which users switch devices and devices are shared across users. This leaves an important question insufficiently examined: whether memory can sustain personalized service **for multiple users across personal and shared devices** while maintaining correct user attribution, up-to-date information, and appropriate access boundaries. To address this gap, we introduce **LILIA**, a benchmark for **L**ong-term memory **I**n mu**L**ti-user, mult**I**-device inter**A**ctions. We first specify who provides information, through which device, and how it is later revised or withdrawn. These events guide dialogue generation and determine reference states and test answers. LILIA comprises **16K multi-turn dialogue sessions in Chinese and English and 4K evaluation instances**. Three complementary tasks assess memory-informed Add, Update, and Delete decisions and their arguments; factual recognition through multiple-choice questions; and open-ended responses involving subject grounding, cross-device use, update and invalidation, and visibility constraints. A unified evaluation harness connects retrieval baselines and memory systems to shared answering and scoring interfaces. We further propose **LILIA-Mem**, which combines joint relation inference with constrained memory revision, preserving conditions and permissions not explicitly changed. Its retrieval process integrates access checks, budget-aware evidence selection, and semantic verification to support context-appropriate responses. The anonymous repository is available at [https://anonymous.4open.science/r/ICLR2027Submission](https://anonymous.4open.science/r/ICLR2027Submission).
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.