acceptodds
Under review as a conference paper at ICLR 2027

Bench2Robust: Scenario-Controlled Recovery for Tool-Using Agents

Abstract

Tool-using LLM agents are typically developed and post-trained in controlled environments with reliable tools, but deployment exposes them to tool failures, API and schema drift, and third-party services outside the developer's control. We introduce BENCH2ROBUST, a framework for scenario-controlled solvability that augments existing benchmarks with episodes requiring retrying the original path, switching to a fallback-equivalent tool, or recognizing that no viable path remains. By controlling recoverability, BENCH2ROBUST separates recovery-policy quality from randomness in whether failures persist or resolve. Across 7 models from 4 families and 10 task/domain slices spanning customer-service and function-calling benchmarks, 69 of 70 model-subset pairs degrade under tool failures, by up to 46.7 percentage points. We propose Bayesian Tool Memory (BTM) for structured runtime recovery context and use reinforcement learning (RL) for learned recovery behavior. On held-out Retail tasks using the benchmark's original tool set, BTM improves robustness by 16.8 percentage points without retraining, while RL retains a 6.3-point gain without inference-time BTM. Combining RL with BTM yields a 20.7-point improvement over the base model while preserving clean performance. These results show that robust tool use benefits from combining structured runtime recovery context with learned recovery behavior.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.