acceptodds
Under review as a conference paper at ICLR 2027

Learning from Success: Online Early Failure Detection for Agentic Systems

Abstract

LLM-based agent systems are increasingly deployed for complex tasks, yet their execution success rates remain limited, making early failure detection critical for raising alerts or intervening in erroneous executions before completion. Existing studies primarily focus on post-hoc failure attribution, which requires complete execution trajectories and forgoes the opportunity to intervene when the task is still running, while the few efforts on early failure detection typically rely on supervised learning from failure trajectories with step-level error annotations. In practical deployments, however, annotating a sufficient number of failure trajectories for supervised failure detection is costly, whereas successful trajectories are naturally accumulated through routine execution. This motivates an important yet underexplored setting: safe-only early failure detection, where an effective monitor identifies failures online by learning exclusively from successful execution trajectories. We introduce a simple framework for this setting, grounded in the insight that successful trajectories provide complementary evidence about both how an execution should evolve and whether the current transition is consistent with previously observed behaviors. Accordingly, our framework combines a safe-dynamics predictor that models the temporal evolution of successful trajectories with a non-parametric safe transition memory that evaluates contextual consistency against historical safe transitions. When a limited number of unsafe trajectories become available, our framework naturally incorporates them as failure memory through causal retrieval and safe-unsafe transition contrast, without additional retraining or fine-tuning. Extensive experiments demonstrate that our framework achieves strong online failure detection performance under the safe-only setting and can further leverage limited failure supervision to approach the performance of fully supervised detectors while requiring fewer annotated failure trajectories.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.