User Queries Say More with Less: Query-Anchored, LLM-Aware Response Selection for Long-Term Conversational Memory
Abstract
In long-term conversational memory, historical user queries and LLM responses exhibit a pronounced asymmetry: user queries are concise and information-dense, whereas LLM responses are substantially longer but often contribute limited new information. We find that historical user queries serve as the primary memory basis, while LLM responses are largely compressible but conditionally useful, with their utility depending on relevance to the current query and the needs of the downstream LLM. Based on these findings, we propose QARS (Query-Anchored Response Selection), a context selection framework that retains historical user queries as memory anchors and selectively incorporates historical LLM responses. It first filters LLM response candidates by their semantic relevance to the current query, thereby narrowing the candidate set; it then performs downstream-LLM-aware selection over the remaining candidates using trained discriminators conditioned on a behavior-based representation of the downstream LLM. Experiments across three long-term conversational memory datasets and six downstream LLMs show that QARS substantially reduces input tokens while maintaining or improving answer accuracy relative to full retrieved contexts, and outperforms existing compression baselines in accuracy. It further constructs LLM-specific contexts and adapts to unseen downstream LLMs using a calibration set without retraining.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.