Proactive Recall for LLMs: Towards Reliable and Timely Task Execution
Abstract
Recent advances in LLM-based dialogue systems have substantially improved their ability in task execution. However, lacking the ability to proactively recall necessary information for pending tasks, they misjudge whether to execute: acting with fabricated values when information is insufficient or staying inactive despite sufficient information already being available in the interaction history. To address this limitation, we formalize Proactive Task Execution (PTE) and introduce PTE-Bench, a benchmark for trustworthy and timely LLM task execution. PTE-Bench encompasses 2,268 long conversations and 10,639 evaluation instances across four conversation scenarios, providing a comprehensive testbed for probing whether an LLM can proactively recall task-relevant information to decide whether to execute, and to identify what is missing when it cannot. We further propose ARDSE, Active Recall through Dynamic State Evolution, which supports proactive recall by dynamically monitoring pending task states through a Progressive Task-State Graph and using evidence-conditioned activation to determine which task to recall in multi-task conversations. Extensive experiments on four LLM backbones and five representative baselines show that ARDSE achieves the best performance, demonstrating the effectiveness of persistent task-state evolution for proactive task execution. Data is available at https://anonymous.4open.science/r/active_recall-65ED.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.