acceptodds
Under review as a conference paper at ICLR 2027

Making Evidence Count: Grounded Diagnosis for Software Agents

Abstract

Diagnosing a software failure requires more than locating relevant code. An agent must connect the failure mechanism to the observed symptoms, identify where and how a repair should occur, and ground these claims in acquired evidence. We make three contributions toward evaluating and supporting this reasoning. First, we introduce FailureSnapshotHub (FSH), an interactive benchmark of 60 real continuous-integration failures from 49 repositories. Agents investigate only failed-side logs, code, tests, configuration, and execution context, while expert-adjudicated references and corroborating developer repair diffs remain hidden. Evaluation separately assesses mechanism, causal direction, repair owner/location, repair direction, completeness, and evidence grounding. Second, we formulate evidence-centric context compilation and present EviCompile, a deterministic method that transforms tool observations into source-linked evidence memory through recoverable text coding, typed evidence construction, cross-observation consolidation, and cost-aware global ordering. EviCompile itself neither generates new observations nor invokes another language model; causal inference remains with the diagnostic agent. Third, we empirically characterize the effects of evidence representation in a sequential Tool Agent, an OpenRCA-style controller–executor agent, and fixed-evidence diagnosis. Case analyses show that agents can locate the relevant code or recognize an immediate mismatch yet infer the wrong mechanism or repair boundary. Compared with original observation histories, EviCompile improves mean semantic scores by 9.72 and 10.42 points in the two interactive settings and by 5.55 points when acquired evidence is held fixed. Local text coding underperforms EviCompile in all three settings; with evidence membership fixed, full ordering also outperforms chronological and relevance-only presentation. Blinded expert reassessment on 40 cases preserves positive mean gains. Together, these results show that evidence organization can improve causal and repair judgments without additional observations, while FSH exposes diagnostic errors that location-based evaluation cannot resolve.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.