Divide, Diagnose, and Compose: Optimizing Agentic Pipelines and Harnesses
Abstract
Building agentic pipelines that achieve high task performance while keeping cost and latency low requires jointly optimizing heterogeneous and coupled design choices, including prompts, topology, model selection, and tool policies. Manually searching this combinatorial design space is expensive and time-consuming, while existing automated optimizers typically operate within a predefined subset of this space, modify the pipeline monolithically, or provide little structured evidence for pipeline changes. To mitigate these challenges, we introduce the Decomposed Credit-Assignment Framework (DeCAF), which adapts the design space to the given benchmark and pipeline, optimizes each dimension separately, and generates reasoning traces for every decision. Given a task specification, benchmark examples, and a seed pipeline, DeCAF identifies relevant optimization dimensions and generates diverse experimental hypotheses for each. Independent coding-agent sessions then optimize each dimension under a restricted edit scope, jointly evaluating performance, inference cost, and latency while recording all hypotheses, edits, and outcomes of each experiment. A final synthesis session combines compatible findings across all dimensions and uses a component-level trajectory evaluator, which derives role-specific rubrics from each agent's instructions, to localize underperforming components and guide cross-dimension revisions. Across a diverse suite of agentic benchmarks, DeCAF improves average task performance by 2.58 percentage points, while reducing inference cost by 4.30% and end-to-end latency by 42.45% compared to the best overall baseline.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.