acceptodds
Under review as a conference paper at ICLR 2027

TotalGuardEval: Benchmark for Evaluating Guardrail Models on Underrepresented Risk Categories

Abstract

Guardrail models—lightweight safety classifiers that filter harmful inputs and outputs of large language models (LLMs)—are advancing rapidly, yet their evaluation remains limited by incomplete taxonomic coverage. Comparing ten widely used guardrail benchmarks and model policy taxonomies, we find critical gaps: Fascism, Religion, and Military conflicts are absent as standalone categories, while Politics receives only partial coverage. We introduce TotalGuardEval, a benchmark targeting these underrepresented categories. It defines a taxonomy of 16 harm categories, each matched with a thematically aligned safe counterpart, and covers both user requests and model responses, enabling measurement of detection performance and over-refusal at the category level. The dataset contains 7.2K examples. We evaluate ten guardrail models, 86M to 20B parameters across three generations, under an original condition, 7 obfuscations, and 88 attack rewrites. The three strongest fall from 0.93–1.00 F1 on original requests to 0.40–0.48 under attack, equally on new and established categories, while one 0.6B model's false positive rate on safe responses spans 0% to 75.5% across categories. Even the most comprehensive prior taxonomy provides at least partial coverage for only 82% of our categories, and no coverage for the three standalone categories we introduce. These results show that current evaluation gives an incomplete picture of guardrail capabilities, and that taxonomically diverse benchmarks are essential for realistic safety assessment.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.