acceptodds
Under review as a conference paper at ICLR 2027

MCPDISCO: SEESAW-MITIGATED TUNING FOR MCP-NATIVE AGENTS

Abstract

The Model Context Protocol (MCP) is rapidly emerging as the substrate of the agentic web: a unified client-server standard that decouples LLM agents from an open-ended, ever-growing space of external tools. As agents become MCP-native, they face a decision-making regime that is qualitatively different from ad hoc function calling—one that requires jointly reasoning over a vast tool ecosystem and producing syntactically precise tool invocations. In this paper, we present \modelname, the first work to perform reinforcement learning (RL) tuning of MCP-native agents at ecosystem scale, spanning 70 heterogeneous MCP servers and 527 tools. We formalize the agent as a Thought-Markov Decision Process (Thought-MDP) that makes internal reasoning a first-class citizen alongside external actions, and we identify a reasoning-tool seesaw effect: under jointly optimized rewards, gains on reasoning tokens systematically trade off against tool-call correctness due to conflicting gradient directions. To resolve this, we propose 1) a disentangled mask reward mechanism that decouples reasoning quality from tool syntax and 2) reasoning-focused gradient surgery that densifies task-relevant reasoning. \modelname achieves higher success with fewer chain-of-thought tokens. On the out-of-distribution MCP-Atlas and GAIA benchmarks, \modelname maintains monotone gains, confirming that the learned behavior generalizes beyond the training server distribution. Model weight and implementation are available. https://anonymous.4open.science/r/MCPDisCo/README.md.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.