TraceSIR: A Multi-Agent Critique Framework for Structured Analysis and Reporting of Agentic Execution Traces
Abstract
Agentic systems augment large language models with external tools and iterative decision making, enabling complex tasks such as deep research, function calling, and coding. However, their long and intricate execution traces make failure diagnosis and root cause analysis extremely challenging. Focusing solely on final task outcomes further discards critical behavioral information required for accurate issue localization. To address these limitations, we introduce **Trace Critique** as the task of producing structured analyses and reports over agentic execution traces, and propose **TraceSIR**, a multi-agent framework that operationalizes this task. TraceSIR coordinates three specialized agents: (1) StructureAgent, which introduces a novel abstraction format, *TraceFormat*, to compress execution traces while preserving essential behavioral information; (2) InsightAgent, which performs fine-grained diagnosis including issue localization, root cause analysis, and optimization suggestions; (3) ReportAgent, which aggregates insights across task instances and generates comprehensive analysis reports. To evaluate TraceSIR, we construct **TraceBench**, covering three real-world agentic scenarios, and introduce *ReportEval*, an evaluation protocol for assessing the quality of analysis reports aligned with industry needs. Experiments show that TraceSIR consistently produces coherent, informative, and actionable reports, significantly outperforming existing approaches across all evaluation dimensions. We further show that TraceSIR provides correct and actionable diagnoses, matching human-annotated root causes, achieving higher scores in a format-independent factual-grounding audit, and enabling substantial recovery of originally failed tasks. Our code is publicly available.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.