OnePred: Next-Query Prediction via Recursive Intent Memory in Multi-Turn Conversations
Abstract
LLM-based conversational systems remain largely reactive: they respond after the user types a query. A key step toward proactive interaction is next-query prediction, which anticipates the user's subsequent query from the preceding dialogue. This task poses an efficiency–quality trade-off: concatenating the full history increases input-token consumption as the conversation grows, while using only the latest turn discards cross-turn context. Our key insight is to track the user's evolving intent trajectory across topics, unresolved needs, and interest shifts, rather than repeatedly re-read the raw history. We propose OnePred, which maintains a recursively updated textual memory as its sole cross-turn context. A two-stage reinforcement learning pipeline first teaches what to predict, then what to preserve, shaping the memory into a prediction-oriented intent chain. We introduce NQP-Bench, spanning three sources of conversations selected for context-grounded predictability, and complement it with deployment-log evaluation without predictability filtering. OnePred improves prediction quality over current-turn and full-history inputs across the evaluated model regimes, with smaller gains over matched learned-summary controls. Its input-token advantage reaches 22× at long turns; accounting for generated memory yields total-token ratios of 4.0–13.5× across dialogue-length groups.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.