acceptodds
Under review as a conference paper at ICLR 2027

MANTRA: Synthesizing SMT-Validated Compliance Benchmarks for Tool-Using LLM Agents

Abstract

Tool-using large language model (LLM) agents are increasingly deployed in settings where their reliable behavior is governed by strict procedural manuals. Ensuring that such agents comply with the rules from these manuals is challenging, as they are typically written for humans in natural language while agent behavior manifests as an execution trace of tool calls. Existing evaluations of LLM agents rely on manually constructed benchmarks or LLM-based judges, which either do not scale or lack reliability for complex, long-horizon manuals. To overcome these limitations, we present MANTRA, a framework that automatically synthesizes machine-checkable compliance benchmarks from a natural-language manual and a tool schema. Since no ground truth exists for translating natural language into formal checks, and even humans are known to be unreliable at such tasks, MANTRA replaces ground truth by redundancy: it independently generates (i) a symbolic world model capturing procedural dependencies, and (ii) a set of trace-level compliance checks for a given task, and validates their consistency using SMT solving. A structured, counterexample-based repair loop resolves inconsistencies, with human review only as a fallback. Importantly, MANTRA supports arbitrary domains and long procedural manuals, and provides a tunable notion of task complexity, so harder tasks can be derived from the same document when agents capabilities improve. Using MANTRA, we build a new benchmark suite with 285 tasks across 6 domains scaling to 50+ page documents with minimal human effort. Empirically, we show that our compliance checks are richer with stronger constraint enforcement compared to existing benchmarks. Additionally, the granularity of the checks can be used for debugging the agents' failure modes. These results demonstrate that combining automated benchmark generation with formally grounded validation methods enables scalable and reliable benchmarking of tool-using agents.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.