Ping: Learning Proactive Agents Over User Trajectories
Abstract
Most existing AI agents are reactive: they wait for users to specify a task, leaving users to notice problems and decide what help to request. On the other hand, a helpful human colleague may offer advice without being asked. In an ideal world, agents should also be able to do the same, offering useful proactive assistance that users might not have considered. Such assistance could help users catch overlooked problems and avoid unnecessary work, without requiring them to ask an agent to check for each possibility. Thus, we introduce Ping, a framework for benchmarking the relevance and timeliness of proactive assistance over user trajectories in software engineering and productivity domains. We derive taxonomies from GitHub comments and meeting transcripts to identify the kinds of proactive assistance agents could offer. We evaluate agents against references that describe the assistance they should proactively offer. Each reference has an assistance window specifying when that assistance should be offered. We use the GitHub taxonomy to extract references for SWE-chat coding trajectories and the meeting taxonomy to categorize existing references in PARE's productivity scenarios. We evaluate five baseline LLMs on 100 SWE-chat test instances and 143 PARE instances. Most references remain unmatched: the highest macro recall among these baselines is 7.0% on SWE-chat and 20.3% on PARE. We also fine-tune agents on user trajectories, using examples of when to offer assistance and when to remain silent. This training increases macro recall on SWE-chat by up to 13.5 percentage points over the corresponding base models. We hope Ping will support the development of proactive agents that can identify useful assistance and offer it at the right time, without waiting for users to ask.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.