acceptodds
Under review as a conference paper at ICLR 2027

SymPolicyBench: Evaluating Symbolic Rule Extraction from LLM-Agent Prompts

Abstract

Enterprise agents are governed by prose instructions, and enforcing that governance independently of the agent’s own reasoning requires re-expressing it as a symbolic policy artifact—structured constraints over tools, parameters, and call counts that another component can read and apply, with a declared naturallanguage form for requirements no structured kind carries faithfully. Systems that synthesize such artifacts are now being built, but the artifact itself is rarely the unit of evaluation: agent-safety benchmarks score a trajectory under a guardrail that is fixed or absent, while synthesis work that does score policies scores them in one domain against one hand-authored target, so no two systems can be compared. We release SymPolicyBench, a corpus of 32 enterprise workflows across 8 domains with 543 golden intents traced to the clauses they derive from and 907 labelled decision points carrying untrusted conversational history. Ground truth is stated as intents rather than rules, so matching is many-to-many and a submission is never penalized for decomposing a requirement differently than the annotator did. Each policy is frozen and scored twice—statically against the intents it should encode, and behaviourally by simulating that exact policy, omissions included, on decision points the scorer sees unlabelled. Across 7 LLM-driven generators under one pinned judge, strict coverage spans 51.7–81.5% while balanced accuracy spans only 7.0 points and inverts at the top, the most faithful configuration obstructing the most legitimate work. Rescoring every frozen submission with a second-family judge and regenerating two configurations establishes which conclusions travel: the fidelity and false-allow orderings reproduce across judges (ρ = 0.82 and ρ = 0.96), and the fidelity–utility trade-off itself holds under a fixed scorer.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.