acceptodds
Under review as a conference paper at ICLR 2027

LMA: Selective Layer-Wise KV Anchoring for Frozen Language Models

Abstract

Safety adaptation should prevent harmful assistance without turning ordinary requests into refusals. A learned safety prefix can strengthen a frozen language model, yet applying it to every request can cause unnecessary refusals. We introduce Layer-wise Memory Anchoring (LMA), which separates learning a safety condition from deciding when to use it. Response targets derived from 28 structured policies train a detachable, 64-slot key–value (KV) prefix at each Transformer layer. Before generation, a lightweight risk router selects the prefix-conditioned or original model path. Its threshold minimizes benign activation subject to at least 99% empirical risk recall on a held-out calibration set. On two frozen 30B Qwen models, selective and always-on use of the same prefix both achieve 99.97% harmful-request protection under Qwen3Guard. Selective activation reduces benign over-refusal from 35.00% to 6.50% on text and from 42.03% to 20.84% on tool contexts, close to the original paths' 5.83% and 21.47%. Controlled target comparisons also favor structured policies over generic action labels and shuffled policies. These results identify activation scope as an effective control for retaining a safety prefix's protection while reducing unnecessary intervention on benign requests.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.