acceptodds
Under review as a conference paper at ICLR 2027

Model Context Agent: Validated Agentic Workflow via Agentic Planning

Abstract

The Model Context Protocol (MCP) has become the standard for connecting large language models to external tools, but it standardizes only how a service is accessed, leaving every client to rediscover how it should be used. That burden falls on a reactive loop that chains tool calls one at a time and judges its own success; on realistic multi-step tasks it is unreliable, costly, and offers no guarantee the goal was met. We argue that the service should ship its own expert. We introduce the Model Context Agent (MCA), an abstraction in which planning, execution, validation are provided by the server rather than delegated to a general-purpose model, and a server-agnostic engine that realizes it: a grounded planner compiles a goal into an ordered tool-call sequence, a deterministic executor replays it with no model in the loop, a verifier confirms the outcome by re-querying the live system instead of trusting a self-report, and a task generator distills verified runs into a knowledge base the planner reuses. We instantiate the same MCA on four production servers spanning 127 tools—GitHub, PostgreSQL, Playwright, and Yahoo Finance—and evaluate it on 125 multi-step tasks under twelve backbones from five providers. Matched against a ReAct loop with the same backbone, tasks, and verifier, it verifies 71% of tasks against 40%, is significantly ahead in 34 of 48 campaigns, and uses up to 12.6 times fewer tokens. On hard suites, the knowledge base an MCA distills from its own verified runs adds 7–12 verified tasks per backbone on PostgreSQL and GitHub and cuts replans by up to 62%.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.