acceptodds
Under review as a conference paper at ICLR 2027

Saturated Screening, Concentrated Cost: What a Full-Stack Ablation Table Reads in a Redundant Refusal Stack

Abstract

Safety teams assemble refusal stacks — released guard models alongside the assistant’s own refusal behaviour — composed so that any one component can block, then decide what the stack is worth by ablating it: switch each component off in turn and see what breaks. We ask what that table measures. We assemble 9 refusal sources (8 released guard models from 5 vendors, plus the assistant’s own refusals), run each guard model at the operating point its vendor ships, and enumerate every one of the 512 subsets over 3937 prompts from 4 suites, so each source’s exact Shapley value sits beside its leave-one-out score with no coalition sampling (attribution figures: main suite). The harmful side is saturated: the full stack passes no harmful prompt on any of the 3 suites that contain harmful prompts, and not one has a sole blocker, a mean of 8.56 [8.46, 8.64] of the 9 sources blocking each StrongREJECT prompt. Switching any single source off therefore changes not one harmful decision at these shipped operating points, so every full-stack leave-one-out number here varies only through over-refusal. A zero full-stack removal marginal need not imply zero coalition-averaged contribution: on the harmful side the exact Shapley values sum to -1.0000 while the leave-one-out entries sum to 0.0000, and ablation sends 5 of the 9 to exactly +0.000 while those same sources carry 71.5% [65.0, 76.5] of under that normalisation, and the loss-minimising subset is drawn from them (WildGuard-7B alone in the unrestricted game, at 11.30 [7.84, 18.51]x lower loss than the full stack, though it passes 20 harmful prompts). So a one-shot rule deleting that column’s zeros removes the sources that carry most of that credit, and the cost is not hypothetical: the union blocks 98.6% [98.0, 99.2] of OR-Bench hard benign prompts. Two things follow cheaply. Calibrating the guard models to a 3% benign false-positive rate on the XSTest-safe pool cuts the union’s benign block rate 3.9x (2.57–7.20x refitted on 4 benign pools) and cuts that same zeroed-mass statistic 4.4x; it also returns 8 harmful prompts to a sole blocker. That is only a start, as 98.4% of the harmful prompts are still blocked by two or more sources, but it shows the saturation is a property of the operating points these guards ship with, not of any-one-blocks composition itself. Nor is the support collapse the 9-way union’s: by coalition size 4, 53.6% of (coalition, member) pairs already read exactly +0.000 on the harmful side (46.7–60.0% with one guard model per vendor). Audit a redundant stack for what it costs: here no single removal changes a harmful decision, so full-stack ablation has no harmful-side support.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.