ProcGround: When Correct Outcomes Mask Ungrounded Decisions in LLM Agents
Abstract
Correct outcomes do not guarantee evidence-grounded decisions. When task priors and narrative cues agree with decisive evidence, agents can follow a plausible default and succeed without consulting that evidence. We therefore introduce **ProcGround**, a benchmark of 215 service tasks across 12 domains. Each task defines an Evidence–Decision Contract, which specifies the evidence required before the first consequential commitment, as well as the subsequent decision and execution requirements. Across 2,365 resolved trajectories from 11 models, Semantic Success is 46.2%, but Contract-Grounded Success is only 24.1%; 47.8% of semantically successful trajectories violate the contract. On valid minimal evidence-flip pairs, initially successful Evidence Bypass cases fail in the flipped world at a rate of 36.4%, compared with 12.5% for Contract-Grounded cases, linking the process distinction to behavioral vulnerability. The phenomenon also appears on external AppBench tasks: 58.5% of trajectories bypass decisive evidence, and 98.5% of these bypasses are masked by correct outcomes. In controlled interventions, a reminder checkpoint improves evidence seeking before commitment for two models, while revealing decisive evidence helps three models recover by the final answer. ProcGround complements outcome evaluation with observable grounding requirements and behavioral tests of sensitivity to decisive evidence.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.