acceptodds
Under review as a conference paper at ICLR 2027

When Do Agent Interventions Help? Matching Interventions to Residual Failures

Abstract

Modern agent systems wrap a language model in skills, execution harnesses, recovery, reflection, and extra inference budget. These interventions, however, are usually evaluated in isolation or under a fixed configuration, obscuring whether an observed gain reflects intrinsic value or merely a match to the failures the surrounding system happens to leave behind. Rather than asking whether an intervention helps on average, we ask when each one helps and what failure it repairs. We run a matched factorial study over nine benchmarks, four models, four skill conditions, four harness configurations, and multiple inference budgets, for nearly 0.9 million agent trajectories. We find that intervention value is conditional rather than intrinsic: no single configuration is optimal everywhere, and the payoff of model, skill, harness, and budget interventions depends strongly on the task and its surrounding setup. This heterogeneity is structured by the residual failure, the failure that remains after the current system has done its best. Trajectory-level failure labels capture part of this structure, but diagnosis does not guarantee repair: even an observable execution failure need not be recoverable. These conditional patterns define capability boundaries and reveal real headroom for selective intervention. Yet the tested lightweight deployable routers fail to beat a fixed Harness policy on quality and cost, despite a far stronger outcome-aware oracle. Agent interventions should be judged by which failure they repair, under which configuration, and at what cost, not by average component gains alone.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.