When Does Snapshot Evaluation Predict Live Microservice RCA?
Abstract
Microservice root-cause analysis (RCA) agents are commonly evaluated on frozen incident snapshots because snapshot evaluation is reproducible and inexpensive. Deployed agents, however, query running systems whose telemetry evolves during diagnosis. It is therefore unclear whether capability estimates obtained from snapshots predict performance in live environments, or which properties of an incident make that extrapolation fail. We introduce a strictly paired evaluation protocol that instantiates each incident under three conditions: Snapshot, with frozen evidence; Live-fixed, with a live backend but the same evidence surface as Snapshot; and Live, with telemetry that grows during investigation. The protocol holds the incident, agent, tool interface, query budget, observation cutoff, and scoring contract fixed, allowing backend interaction to be separated from the value of additional evidence. We further define three machine-checkable validity conditions: evidence sufficiency, execution-identity consistency, and scoring comparability. These conditions must hold before a paired result is attributable. Runs that violate them are classified as undiagnosable or incomparable rather than counted as model failures. On a cross-application benchmark spanning multiple fault mechanisms, we use paired score differences, strict-outcome agreement, and case-level flips to measure Snapshot-to-Live transfer across models, treating independent case families rather than repeated calls as the statistical unit. A complementary failure analysis demonstrates how omitted validity conditions can leave apparently successful evaluations without a defensible comparison. Together, the protocol and benchmark turn the assumption that snapshot evaluation proxies live RCA into a controlled, auditable, and falsifiable empirical question.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.