PLOT: Compiling Policy Documents into Decision Tools for LLM Agents
Abstract
LLM agents must follow domain-specific policies when using tools to automate business workflows. Beyond retrieving policy text, they must combine rules across documents, evaluate conditions and exceptions against the current task state, and decide which actions to take and how to execute them through tools. Repeating this interpretation during execution and across tasks can lead to missed requirements and inconsistent decisions. We introduce PLOT, which enhances agent actions with reusable decision tools constructed on demand. For each request, PLOT explicitly declares the required policy-dependent decisions, such as eligibility or permissible amounts, and resolves them by retrieving previously constructed decision tools or building new ones grounded in relevant policy documents. Each tool takes task-specific facts as inputs and applies reusable policy logic to compute a policy-dependent decision. Tools that pass admission checks on source references and code structure enter a decision tool library, and a separate execution agent calls them to complete the task request. We evaluate PLOT on τ -Knowledge and AutomationBench using GPT-5.2 and DeepSeek-V4-Flash. PLOT outperforms the strongest baseline in every setting, improving passˆ1 on τ -Knowledge by 3.8% and partial credit on AutomationBench by 4.7%, on average across the models. It also uses 11.3% fewer tokens than ACE on average. Ablations on τ -Knowledge suggest benefits from slot declaration, admission checks, and cross-task decision tool reuse. Replacing executable decision tools with cited natural-language rules reduces task success from 36.1% to 20.6%.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.