AgentDriftBench: Measuring Adaptation Dynamics of LLM Agents under Mid-Evaluation Environment Drift
Abstract
Deployed LLM agents operate in environments that change underneath them: API parameters are renamed, required fields appear, tools are deprecated—often mid-task. Static benchmarks cannot see the resulting failure mode: an agent may score well on a fixed suite yet collapse under a schema change and recover slowly, partially, or never. We introduce AgentDriftBench, a benchmark that injects controlled, mid-evaluation environment drift into multi-segment tool-use episodes and measures whether and how agents recover. AgentDriftBench contributes: (1) a drift engine with five drift families applied at segment boundaries, shipped with a treated-vs-control validity protocol (drift halves target-tool hit rate: 0.24 vs. 0.47, z=-3.4); (2) an adaptation-centric metric suite—post-drift success-curve AUC, Recovery Steps, Oscillation Index, and recovery rate—that scores recovery trajectories rather than end states; and (3) a controlled study of memory (scratchpad) and drift-notice prompting. Evaluating six Qwen models (0.5B–7B) on 100 six-segment episodes, we find: (i) a significant scale gradient in absolute adaptation (0.5B AUC 0.207 < 1.5B 0.312, p=0.0001), but normalized recovery is nearly flat—the absolute gap is largely baseline capability; (ii) memory value flips sign across models: scratchpad significantly hurts 0.5B (p=0.004) and helps Coder-3B (p=0.011), with Qwen3-0.6B borderline (p=0.058); (iii) drift notices backfire—conditionally: on 1.5B the notice reduces adaptation (p=0.009), and a budget–quality decomposition shows the harm tracks reduced tool-call budgets, not information content; (iv) drift without an error signal (D4) creates persistent deficits (up to -0.085, p<0.0001) that error-driven adaptation cannot cover. These phenomena are invisible to static benchmarks; AgentDriftBench makes adaptation dynamics a first-class measurement target.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.