GoalBench: Benchmarking goal inference in multi-turn interactions
Abstract
Our objective is to measure the ability of AI systems to infer the goals of human users in naturalistic, multi-turn interactions, a crucial step toward the development of controllable AI systems. However, standardized evaluation remains a challenge due to the difficulty of collecting ground truth data, and scalably assessing model outputs against this ground truth. To address this, we developed a methodology for annotation and evaluation, which uses third-person goal inferences provided by human annotators who viewed a conversation as a reference class for model goal inferences in that same conversation. We introduce GoalBench, which integrates a dataset of 4,731 human-written, validated goals across 776 conversation turns, as well as a dataset of 4,206 human goal pair similarity judgments for calibration of an automated grader, to evaluate goal inference capabilities in real human-AI interactions. We evaluated 15 models on this benchmark; while models generally perform comparably to a human baseline at producing goals contained within the human consensus, and outperform humans at recovering consensus goals, they exhibit systematic differences in the goals they produce. In particular, humans used more hedging language in their descriptions of goals and inferred more general goals relating to user motivations. Our work is a concrete step toward reliable evaluation of goal inference capabilities in realistic interactions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.