System Attention Bias: Zero-Tax Policy Alignment via Inference-Time Logit Modulation
Abstract
Large language models follow system-prompt policies unreliably, and improving compliance usually requires fine-tuning. We introduce System Attention Bias (SAB), a training-free intervention that adds log β to the pre-softmax attention logits wherever a non-system query attends to a system key, multiplying those keys’ unnormalised attention scores by β. SAB adds no parameters, auxiliary model or forward pass, and β can be changed per request on a loaded model. We pair it with a calibration protocol that selects β under a capability budget and evaluate on five instruction-tuned models (0.5B–8B, four families). With β selected on held-out AdvBench splits, SAB gives a lower AdvBench harmful-compliance rate than the safety prompt for Qwen-0.5B and SmolLM2-1.7B (15.3% → 10.1%) at 93.9– 96.6% MMLU retention; the other models are at the floor, transfer to HarmBench is not significant for any model, and TinyLlama-1.1B’s drop from 87.5% to 8.8% is an artefact: it consists of degenerate responses that restate the system prompt, and a control shows that it comes from our implementation also biasing the beginning- of-sequence token and vanishes when only the system block is biased. Compared under the HarmBench classifier with prompt repetition and classifier-free guidance, SAB is not the strongest method: all three lower harmful compliance only by refusing more safe requests (XSTest), and at matched over-refusal they coincide on AdvBench. SAB reaches that operating point in one forward pass (0.2–1.7% overhead per token) where guidance needs two, and, unlike repetition, offers a continuous dial. On role-specific policies (CoSApien), SAB raises Llama-3.1- 8B’s CoSA-Score from 0.423 to 0.457 under greedy decoding, a gain that does not persist under sampling. Over-biasing degrades both safety and capability, so calibration is necessary. The reference implementation uses a dense O(n2 ) mask; the same bias can also be computed inside a fused FlexAttention kernel without materialising it.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.