Shapley-Calibrated Semantic Routing for Black-Box LLM Guardrails: What a Fixed Route Retains
Abstract
Input attribution localizes model behavior to specific prompt components, turning debiasing from untargeted rewriting into targeted intervention. A deployed guardrail, however, must commit in advance to the component it will act on. We study what is lost when prompt-level attribution is compressed into a single fixed routing decision. We introduce Block Prior (BP), which decomposes prompts into four interpretable roles (Subject, Argument, Action, and Setting). BP computes exact four-player Shapley attributions offline under an auditable input-side bias function and aggregates them into one regime-level route applied unchanged to every incoming prompt. Exact attribution costs coalition evaluations per prompt and is tractable both offline and online. The question we answer is how much localization a single stored coordinate can carry. On social-bias regimes from BBQ, the four-role calibration issues a falsifiable prediction about which regimes should depart from a fixed Subject-only policy, and held-out intervention confirms it: +39.2 points of bias reduction for Age and +6.5 for Physical Appearance, with exact coincidence for Gender Identity, where the two policies are identical by construction. The compressed route recovers 77% of the bias reduction obtained by recomputing exact Shapley attribution for every prompt (54.5% versus 70.7% pooled) while improving on the fixed Subject-only policy by 11.2 points, and it preserves high prompt similarity (0.929). We then map where the compression works and where it fails: it is near-insensitive to additive padding (<1% attack success), degrades under out-of-vocabulary lexical substitution (≈40%), and fails completely when bias-bearing evidence is relocated to a non-routed role (100%), a perturbation against which prompt-adaptive attribution remains effective. Cross-dataset and cross-regime studies establish, respectively, ordinal profile agreement on a target-conditioned subset and representational portability of the four-role interface under a different risk specification. Our results quantify the price of a fixed attribution-derived policy
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.