Co-Batching Can Change What Agents Do: Tool Execution and History Effects
Abstract
Tool-using agents turn model outputs into operations that read information, modify records, and complete tasks. Inference servers batch unrelated requests for efficiency, but can these shared execution conditions change an agent’s behavior as well as its output text? We study this question in AppWorld by holding the target request fixed, varying co-batch conditions, executing the resulting actions, and following their consequences and action–observation histories. Repeatable operational forks occur in four of five tested deployment panels. Paired continuations differ on task checks in 36 of 97 evaluable selected states within an eight-decision window, including completed tasks. Restoring common measured database contents still leaves different supplied histories capable of changing subsequent actions. These differences persist across reversed comparison orders, while matched same-history controls agree. On new tasks, average task-check disagreement beyond repeated-condition background is positive but uncertain on both deployments. Deterministic configurations eliminate the tested initiating forks in controlled Qwen and DeepSeek audits; MiniMax shows partial prevention alongside residual and induced differences. The findings show that serving variation can affect executable operations and that database agreement alone is insufficient for behavioral agreement. Evaluating agent reproducibility therefore requires observing serving conditions, executed behavior, and retained interaction history together.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.