Beyond Refuse-or-Comply: Fine-Grained Policy-to-Action Grounding in Safety-Aligned Language Models
Abstract
Safety policies rarely reduce to a binary choice between refusal and compliance: the same request may require unrestricted help, bounded assistance, clarification, refusal, or escalation depending on explicit rules and context. We formalize this problem as fine-grained policy-to-action grounding and introduce PolicyGround, a benchmark with 2,400 controlled counterfactual instances and 1,026 diverse instances. On the Controlled split, five open-weight baselines plus Gemini-3.8-Flash under full-JSON prompting achieve high coarse refusal accuracy (0.880–0.991), yet 12.5–32.8% of coarse-correct outputs still select the wrong fine-grained action. Exact policy citation is also insufficient: across full-JSON Controlled and Diverse evaluations, 5.0–34.3% of outputs that cite exactly the required rules still choose the wrong action. To separate marginal lexical cues from request-policy binding, we report request-only and policy-only audits and introduce a Latin-square crossed diagnostic in which every request and policy is paired once with each action. Additive TF–IDF remains at chance on this diagnostic, while language models reach 0.788–1.000; the strongest models saturate explicit binding but remain imperfect on less templated action boundaries. Family-disjoint SFT-LoRA improves held-out and Diverse action accuracy, but increases unsafe-compliance errors on Diverse. PolicyGround separates coarse correctness, citation compliance, explicit binding, and action calibration, while leaving response-level semantic safety to future work.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.