acceptodds
Under review as a conference paper at ICLR 2027

Failure Attribution with Stepwise Uncertainty in Multi Agent Systems

Abstract

Failure attribution in agent systems aims to identify the steps that critically contribute to task failure. Existing approaches either rely on large language models (LLMs) as judges, incurring substantial inference costs, or require additional data and training to build dedicated attribution models. We propose a training-free method for failure attribution that identifies critical failure steps using the outputs of a fixed, lightweight observer model. We find that decisive failure steps leave a measurable signature in the observer's continuous uncertainty. Conventional binary readouts that collapse predictions into binary outcomes discard relative uncertainty across steps, but the proposed method directly uses the observer model's pre-threshold uncertainty as a continuous attribution score, identifying the first step whose uncertainty enters a high-uncertainty regime as the critical failure step. It requires no modification to the observer model or its inference procedure, and introduces no additional training data, step-level failure labels, or auxiliary evaluation models. Recovering the continuous information enables fine-grained, step-level failure attribution at minimal additional cost. Experiments on the Who & When benchmark demonstrate that this attribution signal persists across multiple unsupervised localization rules, and that the proposed method outperforms closed-model LLM-as-a-Judge baselines by up to 10.3%p, while avoiding $2.11-$14.44 in frontier-model API costs per corpus without requiring any additional model calls.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.