PAIRBench: Complete Response-Function Coverage for Auditing Policy-Conditioned Safety Guards
Abstract
Policy-conditioned safety guards decide whether a request violates a policy supplied at runtime. Policy revisions require correct updates on affected requests and stable decisions elsewhere. High cell-level scores can coexist with incomplete recovery of joint policy–scenario patterns, while selected relational tests leave other patterns unmeasured. We introduce PAIRBENCH, a benchmark covering all 16 deterministic binary response functions formed by two policies and two scenarios. Its human-reviewed panel balances 160 quartets; averages exact recovery equally across functions, and matched policy interventions test updates and stability with scenarios fixed. DynaGuard and SingGuard achieve cell accuracies of 0.969 and 0.986 but scores of 0.888 and 0.944. Complete stratification yields lower root-mean-square error for population-level exact recovery than the tested reduced panels at a matched 80-quartet budget. Four guard families meet prospectively frozen selective-update criteria, and a separate policy-replacement control confirms behavioral dependence on the supplied policy in three tested systems. On a harder document-rich panel, LPG and SingGuard show lower selective updating and transition recovery. The benchmark connects complete local function coverage with controlled tests of policy response, complementing per-decision scores with an explicit population audit.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.