APEX: Adaptive Progressive Examination of Agentic Workflows
Abstract
Large Language Model (LLM)-based agentic workflows have shown substantial gains over single agents on complex tasks, but they are harder to investigate when they underperform, since a weakness can lie in any aspect of any of their units: the agents, the communication strategy among them, the memory they share, and the tools through which they act. Existing failure attribution analyzes the trajectory of a run that has already failed and returns a single responsible agent and step, so nothing is learned until a failure has occurred, failures are attributed to agents alone, and the aspect in which the responsible unit is weak remains unidentified. We formulate unit-level weakness localization, which identifies the weak units of a workflow and the aspects in which they are weak before they fail in deployment. We present (1) APEX, which adaptively and progressively explores a workflow's execution trajectories, running probes that target the assertions of a dynamically generated rubric tree about which it is statistically least confident; (2) AXE, an execution engine that records what every unit was given and changed, and can perturb any unit in isolation; and (3) AXEBench, a labeled benchmark of perturbed workflows generated by AXE. On AXEBench, APEX correctly localizes 90.5% of injected weaknesses in agents, against 57.1% for the state of the art, and 51.5% across all units of the workflow, against 14.5%, as prior methods can only target agents; it identifies the weak aspect in 79.0% of the weaknesses it localizes. Guided by APEX, an LLM instructed to repair the workflow recovers more than twice as much of the lost performance as it does without guidance, and human experts validate both the verdicts APEX relies on and the benchmark it is evaluated on. Code and data are publicly available.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.