MISSING MASS, MISPLACED CONFIDENCE: CALIBRATING GENERATIVE SAFETY GUARDS
Abstract
Generative safety guards produce scores that downstream systems threshold, rank, and audit as probabilities. A commonly used score renormalizes the guard's next-token distribution over the safe and unsafe label tokens, discarding their total probability mass, which we call verdict mass, . For a fixed input prompt, the resulting conditional distribution differs from the full output distribution by exactly in total variation, and the guard's first-token log loss separates into mass and verdict-choice terms. Across three Llama Guard checkpoints, the released prompt templates concentrate mass near one, while removing their safety instruction exposes substantial variation across inputs. For the same input texts, safety-assessment instructions yield higher verdict mass than length-matched general-task instructions, and attack rewrites that the guard misses flip its verdict while largely preserving this mass. We combine the instruction-free mass with the native logit margin in verdict-mass calibration (VMC), a closed-form correction requiring no guard retraining, in a parameter-free form and an extension with one fitted temperature. On a held-out mixture of benign, harmful, and attack inputs, the extension reduces negative log-likelihood by approximately 9–38% relative to temperature scaling and improves Brier score and expected calibration error across all three checkpoints; it also improves all three metrics on ShieldGemma 9B.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.