Interactive Trajectory Environments for Long-Horizon Agent Diagnosis
Abstract
Diagnosing agentic systems requires identifying where an execution went wrong, which becomes increasingly challenging for long-horizon executions as the diagnostic space grows and relevant evidence becomes dispersed across distant but dependent steps. Existing approaches primarily focus on designing the diagnosis agent or procedure, while treating the execution trajectory as passive context. We instead formulate diagnosis as agent–environment interaction, jointly designing the diagnosis agent and the environment through which it investigates an execution. We introduce TIDE (Trajectory-to-Interactive Diagnosis Environment), which transforms a failed execution trajectory into an interactive environment for diagnosis. Specifically, we construct a dependency graph over execution steps and hierarchically organize it into execution processes; together, these structures define an environment in which the agent can navigate across levels of abstraction and inspect dependency-related evidence. An Agent-as-a-Judge adaptively interacts with this environment to explore the diagnostic space, acquire relevant evidence, and localize the erroneous step. Experiments on two benchmarks show that TIDE consistently improves error localization across agentic systems and judge models, reducing mean localization distance by up to 55% on long-horizon trajectories. Further analyses show that dependency-aware exploration is particularly beneficial when diagnostically related steps are distant in execution order.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.