PrismEval: A Benchmark for Fine-Grained LLM Compliance Evaluation
Abstract
LLM compliance spans privacy, copyright, and content safety, and the behavior under evaluation changes with the applicable jurisdiction, language, cultural context, and request form. We introduce PrismEval, a fine-grained benchmark that constructs each dimension from its own source evidence: protected data and subject types for Privacy, protected works and reproduction requests for Copyright, and legal, policy, official, and cultural materials for Content Safety. PrismEval covers five languages and strengthens the resulting test cases through 15 jailbreak methods, five controlled transformations, Dialectal Transformation, and Local Model Rewriting. The complete Jailbreak Enhancement collection contains 32,238 cases, while the multilingual Privacy and Content Safety suite contains 37,433 cases. Expert audits report usable rates of 100% for Content Safety, 100% for Copyright, and 98% for Privacy; both local transformation procedures also produce response differences exceeding three times their within seed baselines. Across 24 models, mean violation rates range from 85.9% to 96.5% across the reported risks, while the lowest overall violation rate remains 78.6%. Five Test Case Optimization Procedures reduce evaluation volume by 40.17% while retaining the delivered model performance characterization. PrismEval exposes large, risk-specific compliance failures across dimensions, languages, and localized requirements.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.