EvoMCP: Co-Evolving MCP Agents and Training Data
Abstract
Supervised fine-tuning of Model Context Protocol (MCP) agents typically relies on static training data that does not adapt to the student’s evolving capabilities. We propose EvoMCP, an agent–data mutual evolution framework that evolves complete task–trajectory pairs through a closed loop of task evolution, trajectory evolution, and supervised fine-tuning. EvoMCP uses student rollout feedback to assess established capabilities and capability gaps, guiding the selection of learning targets and compatible verified seed tasks. Task evolution translates these targets into executable tasks by adapting task instructions, observable environment states, or both, while maintaining consistency among instructions, environments, and verifiers. Trajectory evolution uses reusable skills to recover failed teacher rollouts and refines successful demonstrations for efficiency. Accepted demonstrations train the next student, while their associated tasks become seeds for subsequent evolution. Fresh feedback from the updated student then guides the next round, allowing the training distribution to evolve alongside the agent’s capabilities. Experiments on Toolathlon, MCPMark, and MCP-Atlas show consistent improvements across three evolution rounds. Starting from Qwen3.6-27B, EvoMCP achieves absolute gains of 16.5, 10.2, and 12.6 percentage points on the three benchmarks, respectively.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.