acceptodds
Under review as a conference paper at ICLR 2027

Echoverse-MCP: Stateful Tool Environments for Agent Learning

Abstract

Reinforcement learning for tool-using agents requires environments that maintain state, enforce application rules, and expose the consequences of actions. Building these environments by hand limits the range of workflows available for training. We present Echoverse-MCP, an LLM-assisted pipeline that combines public OpenAPI specifications with reusable domain models to construct isolated, stateful Model Context Protocol environments. Each environment pairs typed tools with a relational database, seeded records, and explicit state transitions. The pipeline generates multi-step tasks by executing reference programs and retaining their tool traces and database changes as evidence. The resulting resource contains 121 environments and 5,916 tools, with 19,041 tasks across 45 environments. Reference solutions average 7.15 calls and 6.24 distinct tools. A separate 600-task evaluation suite covers scripted-user and simulated-user interaction. Mixed GRPO post-training of Qwen3.5-4B raises reported BFCL accuracy by 5.6 percentage points and mean three-domain pass@1 by 3.2 points. The results support the use of generated executable environments for agent learning while showing that aggregate gains can coexist with regressions in individual capabilities.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.