IMPACT: Measuring Conditional Workflow Responses to Single-Span Model Substitution in Stateful LLM Agents
Abstract
Selectively replacing individual language-model invocations with smaller models can reduce inference cost in LLM agents. Yet an agent call is not an independent request: its output can alter subsequent actions, state transitions, observations, and computation. Local model quality and invocation cost therefore provide only a partial view of a substitution, whose effects may propagate through the rest of the workflow. We present IMPACT, a workload-agnostic intervention methodology for measuring conditional workflow responses to single-span model substitution. Starting from a common restored pre-invocation state, IMPACT uses a four-condition protocol to separate replay fidelity, continuation variation, baseline-model resampling, and candidate substitution. It estimates substitution effects from repeated matched baseline continuations, treating the recorded trajectory as one realization of the baseline workflow. IMPACT characterizes each intervention through local behavior, downstream propagation, final task quality and collateral state outcomes, and resource use, and relates these responses to conditions available before the target invocation. We instantiate IMPACT across six cohorts spanning AppWorld, LoCoMo, and SWE-bench, four agent implementations, and 17,100 live continuations. Across these populations, baseline executions vary, local agreement and final quality often diverge, and resource savings do not consistently align with task outcomes. The resulting conditional response atlas summarizes how substitution effects vary with execution context and population and provides evidence for offloading decisions under explicit quality, risk, and resource objectives.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.