acceptodds
Under review as a conference paper at ICLR 2027

TracePrism: Macro-to-Micro Diagnosis for Long-Horizon Targeted Agent Optimization

Abstract

Optimizing long-horizon LLM agents from execution traces requires deciding which failures to analyze, where they originate, and which components to revise. Large trace collections contain recurring failures, while long executions can obscure the steps and component interactions underlying an observed failure. We present TracePrism, a macro-to-micro framework for budget-aware bottleneck localization and agent optimization. At the macro level, TracePrism groups failed trajectories by execution structure and failure information, then selects representative cases under a fixed trajectory budget for detailed diagnosis. At the micro level, it follows data and control dependencies backward from failure manifestations to construct diagnostic slices, localize plausible root-cause steps, and identify editable intervention targets. These diagnoses guide component-specific revisions that are checked against component responsibilities and applied jointly, leaving untargeted specifications unchanged. Across VeruSAGE-Bench, WebArena, and HotpotQA, TracePrism achieves the highest mean task performance among the evaluated methods. On VeruSAGE-Bench, it improves the success rate of an expert-designed agent from 46.2% to 58.8%, a gain of 12.6 percentage points. Scaling experiments on this benchmark further show a favorable trade-off between task performance and optimizer-side LLM cost as the failure corpus grows. The code is available at https://anonymous.4open.science/r/TracePrism-8CC3.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.