acceptodds
Under review as a conference paper at ICLR 2027

Repair Reach: SafeLoRA's Gate Is a Public Scalar, and the Attacker Reads It Too

Abstract

Post-hoc safety repairs promise to undo harmful fine-tuning cheaply, each through a set fixed before the attack is seen, such as SafeLoRA's alignment matrix from public checkpoints, the row space of gradient patching, RESTA's direction or a coordinate mask's support. We call that set the repair's reach. SafeLoRA's gate is public too: the scalar that gates its threshold mode is a closed form in the alignment matrix's singular values, weighted by the update's energy on each, a value the attacker computes before training and steers. Constraining the adapter's output factor into the leading singular subspace (one thin SVD per matrix, a projection per step, with data, objective, optimiser, rank and locus unchanged) makes the gate read the attack as already aligned, and on three instruct models the repaired model still complies with 52.5–62.9% of harmful requests, against 57.9–62.1% unrepaired, while the defence takes 24–59 points off every ordinary family mean; no threshold removes the attack while leaving the customer half their gain on their worst task. A second adversary, constrained out of the public safety span, walks past our coordinate mask at rank 16 and the 20% budget we set; 40% catches it on Qwen2.5-1.5B, and on Qwen2.5-7B only gate and mask together do. Read back as a detector the gate fails, on one model, to two statistics an attacker buys off separately. The study covers four models, 3,086 cells and 367,138 judged generations. Code and per-cell records will be made public upon acceptance.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.