acceptodds
Under review as a conference paper at ICLR 2027

PrismEval: A Benchmark for Fine-Grained LLM Compliance Evaluation

Abstract

LLM compliance spans privacy, copyright, and content safety, and the behavior under evaluation changes with the applicable jurisdiction, language, cultural context, and request form. We introduce PrismEval, a fine-grained benchmark that constructs each dimension from its own source evidence: protected data and subject types for Privacy, protected works and reproduction requests for Copyright, and legal, policy, official, and cultural materials for Content Safety. PrismEval covers five languages and strengthens the resulting test cases through 15 jailbreak methods, five controlled transformations, Dialectal Transformation, and Local Model Rewriting. The complete Jailbreak Enhancement collection contains 32,238 cases, while the multilingual Privacy and Content Safety suite contains 37,433 cases. Expert audits report usable rates of 100% for Content Safety, 100% for Copyright, and 98% for Privacy; both local transformation procedures also produce response differences exceeding three times their within seed baselines. Across 24 models, mean violation rates range from 85.9% to 96.5% across the reported risks, while the lowest overall violation rate remains 78.6%. Five Test Case Optimization Procedures reduce evaluation volume by 40.17% while retaining the delivered model performance characterization. PrismEval exposes large, risk-specific compliance failures across dimensions, languages, and localized requirements.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.