Action Admission and Constraint Evolution: Governing LLM Agent Actions and Learning from Failures
Abstract
LLM-based agents propose actions with consequential effects, some of which violate system constraints. We introduce Action Admission and Constraint Evolution (AACE), a model-agnostic framework that separates action proposal from execution. An Action Admission Controller checks each proposal against predefined constraints and a Constraint Bank. Verified Constraint Evolution generates constraints from observed failures and admits them only after empirical verification and a Candidate Curation Gate. We evaluate AACE in negotiations adapted from AgenticPay and procurement scenarios adapted from RetailBench. In AgenticPay, three of 26 candidate families pass verification and curation, all concerning privacy. On 60 held-out profiles across 540 episodes, the validated Bank reduces attack violations from 57.0% to 31.1% with no observed benign false blocks; admitting all candidates blocks 13.3% of benign episodes. With a Claude executor, the same frozen Bank reduces attack violations from 52.8% to 25.0%. Static admission with deterministic recovery raises valid success from 12% to 96% in the GPT-5.4-mini experiment. In a matched comparison with ToolGuard and a one-revision host, both methods have zero observed executed violations and show no significant difference in valid success (88% versus 92%); AACE has higher mean seller reward and lower latency in this setup. In RetailBench, a verified constraint prevents duplicate procurement in all 12 held-out risky scenarios and blocks none of the 12 benign scenarios. These results support verified constraint updates for specific failure patterns.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.