Mind the Tool Failures: Achieving Synergistic Tool Gains for Medical Agents
Abstract
For medical AI agents, disagreement among external tools can signal both a risk of error and the potential for better decisions: tools that fail on different instances may together outperform even the strongest individual tool. We define tool synergy as the ability to realize these complementary gains and characterize its potential through the Single-Oracle risk gap, which measures the performance gap between the best fixed single tool and an ideal instance-wise tool-use policy. Yet instance-level adaptation alone does not guarantee synergy, as standard outcome-based reinforcement learning and conventional aggregation methods do not explicitly exploit disagreement to learn when and how tools can complement one another. To address this challenge, we propose Collaborative Synergy Reinforcement Learning (CSRL), a GRPO-based framework that turns tool disagreement into an explicit learning signal for synergy. CSRL combines a Brier reward for probabilistic risk minimization, an override reward for disagreement-aware conflict resolution, and Entropy-Guided Sampling that prioritizes high-disagreement cases with stronger signals for learning complementary tool use. Across two tasks and seven medical benchmarks, CSRL consistently outperforms a broad range of baselines, with its advantage rising from 0.4% under tool consensus to 6.4% and 7.2% under moderate and high disagreement, respectively, highlighting its ability to turn tool disagreement into complementary gains.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.