acceptodds
Under review as a conference paper at ICLR 2027

Cirdan: Structured Deep Investigation for Long-Horizon Code Diagnosis

Abstract

Diagnosing distributed system failures requires long-horizon investigation: LLM agents must trace causality across components, connect observed symptoms to responsible code, and distinguish root causes from downstream symptoms. Through an empirical study, we observe that agents often settle on plausible explanations but leave key causal steps unexplored. We introduce Cirdan, a framework for structured investigation of such failures. Cirdan performs broad repository survey, triages competing hypotheses, and explores promising explanations in a tree that supports refinement and backtracking. Its context management retains detailed evidence along each branch and summaries of alternatives, while preserving shared prefixes for prompt-cache reuse. Tool-call context injection brings in structural and architectural information to help the agent look beyond the currently inspected subsystem. We evaluate Cirdan against four agent frameworks across five models on 34 real-world bugs from 11 distributed systems. Cirdan improves diagnosis accuracy with four of the five models. It raises block coverage by 6.3–22.7 percentage points with GPT-5.2-Codex and by 7.4 and 3.6 points with Opus 5 and Gemini-3.8-Flash. With GPT-6-Astra, it maintains coverage comparable to mini-SWE-agent's while reporting half as many code blocks, raising precision from 42.2% to 61.4%. Our results show that keeping competing explanations and their evidence available, while leaving investigation decisions to the model, improves code diagnosis with most models, and that the harness's budgets and guidance must match how each model reasons.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.