Agentic RCA under Partial Observability: From Root Observation to Diagnosis
Abstract
LLM-based root-cause analysis (RCA) agents iteratively inspect metrics, logs, and traces to identify the source of an incident, but fault propagation and concurrent activity can make non-root components plausible competing explanations. Final-answer accuracy alone cannot reveal whether telemetry from the root-cause component was retrieved, whether the root was considered or favored, or whether it reached the final diagnosis. We analyze 335 OpenRCA queries by aligning incident-side signal conditions with retrieved telemetry, expressed diagnostic states, and final diagnoses. Later root signals and especially stronger non-root competition accompany poorer diagnosis, whereas broader anomalous activity does not show the same pattern. Root retrieval shows no consistent association with non-root competition. After observing the root, however, agents facing stronger competition are less likely to favor it or select it in the final diagnosis. We further evaluate three targeted agent designs on 22 queries. Although these designs improve component accuracy, residual failures show that additional structure can leave competition unresolved. These findings motivate cross-hypothesis discrimination—using evidence to distinguish among causal hypotheses—as a design target for agentic RCA.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.