acceptodds
Under review as a conference paper at ICLR 2027

Where Documentation Ends: Source-Derived Behavioral Specifications for LLM Tool Simulation

Abstract

A common way to evaluate, and increasingly to train, LLM agents is to replace the tool backends in their sandboxes with an LLM that simulates each tool response, because real backends are costly, unstable, or unsafe to call at scale. These simulators are usually conditioned on API documentation, while real backends also enforce rules that no document states. We study this documentation gap on three author-built tool servers with 38 frozen hidden rules, 31 of which never fire in any demonstration. Doc-only failures concentrate on rules that contradict a documentation-shaped prior (17.2% vs. 69.6% accuracy), reference-free plausibility judges do not detect them (Youden J <= 0.04), and they change what agents do under real-backend replay. Comparing channels for teaching a simulator the never-triggered rules, specifications derived offline from source reach 80.3% first-encounter accuracy, against 22.7% for documentation and 59.1% for 64-call active probing; a generic summarization prompt does as well, so the gain comes from the channel, not the compilation method. Downstream, source-derived specifications repaired both focal doc-only failures, with success depending on whether a rule fired rather than on its exact value. On three public MCP servers, undocumented behavior, judge blindness, and false-success outcome changes all recur.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.