When Safety Judges Agree Too Much: Correlation Breaks Naive Fusion and Caps Its Gains
Abstract
Deployed LLM systems increasingly run several safety guards and combine their verdicts, either by multiplying their evidence or by escalating from a cheap guard to expensive ones. We give a statistical account of how the guards' agreement, together with their individual strength, governs what such combinations can and cannot do. Across seven public guards from four families and three human-labelled benchmarks, guards are strongly correlated, and on WildGuardTest and BeaverTails all of them fail together on 12–13% of unsafe responses, against at most 0.03% under independence. We prove that this correlation silently breaks naive fusion: the product of per-guard likelihood ratios is no longer an e-value, correlation inflates its null mean and log-variance by closed-form factors, and its false-alarm rate tends to a positive floor as more equicorrelated guards are added, for any fixed nominal level; empirically, the common escalation cascade exceeds a nominal 5% false-alarm rate in 6 of 7 settings, reaching 26%. Correlation can also cap the gain of fusion, and within a Gaussian-copula model does so quantitatively: a single quantity, the information the remaining guards add to a cheap one, bounds both what any combination rule, including adaptive cascades of arbitrary cost, can gain and what calling only the cheap guard loses. On real scores we use the model for gap sizes and rankings, not as a certified ceiling. We complement this with a cascade construction whose false-alarm rate is controlled in finite samples for any calling rule fixed before calibration and any dependence, using about 100 target-domain safe examples exchangeable with future safe traffic. Across seven in-domain and shifted settings, a static subset chosen on source data, at about a tenth of the full ensemble's parameter memory, retains nearly all of its detection at a calibrated 5% false-alarm rate, though less at stricter levels. An adaptive dependency-aware cascade does not consistently improve on it, and, for these seven guards and under the same target calibration, the largest guard alone is never more than 0.001 below a calibrated fusion of all seven in detection, in point estimates.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.