Ground-Truth Lifetime Analysis: Localizing Failures in Agentic Root Cause Analysis
Abstract
Root cause analysis (RCA) agents are judged by whether their final diagnosis is correct, which tells a developer that a run failed but not where. Recent work focuses on identifying hallucination and execution errors in reasoning traces. However, we find that many failures are well-formed on their own terms: every tool call succeeds and every claim is grounded in what the agent observed, yet the true cause was already lost somewhere earlier in the run. We therefore introduce Ground-Truth Lifetime Analysis (GLA), which localizes a failed diagnosis to the stage at which the ground truth was lost — never observed, observed but never raised as a candidate, or raised but rejected. The underlying philosophy is to make the agent's hypotheses auditable, so GLA instruments the agent loop to record its operations on a candidate set at every step, and attributes each failure by tracking the true cause's membership in that set, with no LLM judge. Measuring four frontier models on microservice systems, we find that one stage dominates: up to 83% of failed runs observed the relevant telemetry but never formed the true cause as a candidate, while the observation gap reaches 42% in the weakest model-system pairing. GLA further exposes how agents respond to interventions, in terms that end-to-end metrics cannot capture: the variation across models, the substantial migration of runs between stages behind an unchanged total, and the bottleneck that shifts rather than clears. Our intervention improves accuracy by up to 13.03%.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.