Analyzing Agentic benchmarks under subgoal interactions
Abstract
Tool-calling is a core function of LLM agents, and a plethora of benchmarks attempt to evaluate it. These benchmarks divide into stateful benchmarks, where tool calls operate on shared state such as a database or environment, and stateless benchmarks, where a composition of tools computes or retrieves a value. Stateful benchmarks such as TauBench, AppWorld, and AgentBench turn out to be ergodic, or are ergodified by the harnesses agents run in, and so mask a failure mode studied extensively in the AI planning literature: negative subgoal interaction. In ergodic environments, negatively interacting subgoals increase the cost of achieving one another; in non-ergodic environments, they can make the goal unreachable entirely. While AI planning has studied this concept at length, no study has examined how it affects agentic planning. We introduce PlanBench-Gym, an interactive, tool-calling extension of PlanBench that evaluates LLM agents on PDDL Blocksworld instances constructed along a three-way taxonomy of no-interaction, positive-interaction, and negative-interaction goal structure. Even though these instances remain ergodic under our construction, we find that state-of-the-art reasoning models struggle with negative interactions, though the degree varies: closed source models such as GPT-5.2 fails outright on negative-interaction tasks under most agent scaffolds, while Claude Sonnet 5 completes them but at a measurable efficiency cost. Additional experiments on open-source models mirror the finding with Claude Sonnet 5, where Qwen3.6, GLM5.2, and Gemma4 are able to consistently solve tasks, but do so at a measurable efficiency cost. We further find that prompting mode, whether an agent observes its full interaction history or only the current state, has a substantial effect on completion, independent of interaction class.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.