Flexible Agent Alignment with Goal Inference from Open-Ended Dialog
Abstract
We introduce Open-Universe Assistance Games (OU-AGs), a formal framework extending assistance games to LLM-based agents. Effective assistance requires reasoning over human preferences that are unbounded, underspecified, and evolving. Existing assistance game formulations assume fixed, predefined preferences, an assumption that breaks down in open-ended dialogue where goals are revised incrementally and expressed in natural language. Grounded in cognitive science accounts of preference construction, we represent human preferences as a dynamically updated distribution over discrete natural-language goals. To operationalize OU-AGs, we introduce GOOD (GOals from Open-ended Dialogue), a data-efficient online method that extracts and ranks candidate goals during interaction, using LLM-simulated users to perform probabilistic inference over goal hypotheses, allowing for interpretable, uncertainty-aware preference representations without large offline datasets. Across text domains in grocery shopping, household robotics (AI2-THOR), and coding, GOOD matches or exceeds baselines without explicit goal tracking in alignment with user intent, with the largest gains in robotics and coding, while producing interpretable goal updates justified by the dialogue. On existing multi-turn preference benchmarks, where the agent's choices have limited influence on how the interaction unfolds, explicit goal tracking helps less, suggesting that its value lies in letting inferred goals shape the agent's actions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.