acceptodds
Under review as a conference paper at ICLR 2027

CATO: Contrastive Attribution for Targeted Optimization of Multi-Component LLM Agents

Abstract

Multi-component LLM agents expose several editable components, such as system prompts, tool descriptions, and experience banks. Yet when an agent fails, its binary task outcome does not indicate which component to edit. Single-component optimizers cannot reach defects elsewhere, and unguided joint search can spend its budget on unrelated components. We formulate this component-selection problem as structural attribution, which assigns a failure to the agent's editable components rather than to trajectory steps. A component is an update target of a failure when a repair of the failure needs to edit it. We propose CATO, which contrasts failures with successful runs of the same tasks under a moderately different configuration, scores how strongly this evidence implicates each component, and writes a textual gradient describing the fix. CATO then proposes more single-component edits for higher-scoring components and keeps the best configurations on a held-out validation set, so multi-component repairs build up one edit at a time. In controlled fault injection, CATO identifies the faulty component in of cases, points more often than trajectory-only diagnosis, and its edits repair of these failures while preserving of solved tasks. Across AppWorld, WebArena, and BBH with three backbones, CATO outperforms all six baselines in every setting, leads the strongest by up to points, and uses fewer training tokens than editing every component on every failure.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.