AgentLTL: A Trace-Verification Framework for Measuring, Enforcing, and Training Procedural Compliance in Tool-Using LLM Agents
Abstract
Tool-using LLM agents are typically evaluated by final-answer correctness or by LLM judges. Neither captures how an answer was produced, yet in safety-critical settings the procedure is itself part of correctness. In this paper, we introduce AgentLTL, a specification language derived from first-order linear temporal logic over finite traces (FO-LTL), which expresses procedural rules over agent traces and yields a deterministic, judge-free compliance score. A single specification serves two purposes. In harnessing, constraints either score completed traces offline or gate tool calls online, by checking each prefix before execution. In finetuning, the same compliance signal acts as a dense reward. We evaluate both on our novel dedicated benchmark ProcFlow covering ordering, branching, iteration, and grounding, and on other several established tool-calling benchmarks. Online harnessing raises procedural compliance across models without degrading task performance, and finetuning with compliance-aware rewards improves held-out performance on ProcFlow while also increasing accuracy over the base model on every external benchmark we evaluate. Together, these results show that compliance is a transferable measure of procedural structure: it generalizes beyond the procedures seen during training and can be leveraged directly for efficient runtime enforcement.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.