Bench2Robust: Scenario-Controlled Recovery for Tool-Using Agents
Abstract
Tool-using LLM agents are typically developed and post-trained in controlled environments with reliable tools, but deployment exposes them to tool failures, API and schema drift, and third-party services outside the developer's control. We introduce BENCH2ROBUST, a framework for scenario-controlled solvability that augments existing benchmarks with episodes requiring retrying the original path, switching to a fallback-equivalent tool, or recognizing that no viable path remains. By controlling recoverability, BENCH2ROBUST separates recovery-policy quality from randomness in whether failures persist or resolve. Across 7 models from 4 families and 10 task/domain slices spanning customer-service and function-calling benchmarks, 69 of 70 model-subset pairs degrade under tool failures, by up to 46.7 percentage points. We propose Bayesian Tool Memory (BTM) for structured runtime recovery context and use reinforcement learning (RL) for learned recovery behavior. On held-out Retail tasks using the benchmark's original tool set, BTM improves robustness by 16.8 percentage points without retraining, while RL retains a 6.3-point gain without inference-time BTM. Combining RL with BTM yields a 20.7-point improvement over the base model while preserving clean performance. These results show that robust tool use benefits from combining structured runtime recovery context with learned recovery behavior.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.