acceptodds
Under review as a conference paper at ICLR 2027

Groundworks: Agents Can Complete Tasks Yet Fail Under Plausible Perturbations In Multi-Turn Interactions

Abstract

Agents carry out tasks through evolving conversations and tool use, yet success in a fixed scenario does not guarantee whether they can handle real-world challenges such as revised requirements or unexpected tool results. Mishandling these can lead to hallucinations, such as misstating the user’s current requirements or claiming success without tool evidence. Thus, task completion alone cannot establish whether their claims are grounded in the evidence available at each turn. We introduce **GroundWorks**, an adaptable, end-to-end framework that constructs paired multi-turn evaluations with verifiable target outcomes. Its 30 reusable perturbation operators model plausible changes to an interaction or environment and target specific hallucination risks. For each operator, GroundWorks creates an original, solvable scenario and a matched perturbed scenario, specifying the relevant evidence and criteria for appropriate handling before execution. It then simulates both interactions and verifies agent behavior against the predefined criteria and recorded evidence, distinguishing successful handling, targeted hallucination, and other failures. Across five agents and 1,040 paired evaluations, agents fail on 47% of perturbed scenarios whose originals they complete, averaged across agents: 25% involve targeted hallucination and 22% other failures. Changes to tool capabilities and action outcomes yield the highest conditional failure rate (71%). We further apply GroundWorks to patient-facing healthcare tasks, demonstrating its adaptability to a new setting.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.