acceptodds
Under review as a conference paper at ICLR 2027

When Does In-House Adaptation Cost Safety? A Capacity-Rank-Recipe Map

Abstract

Fine-tuning an assistant on an in-house security corpus looks like the benign case: domain knowledge padded with ordinary alignment data. It is not. We fine-tune Qwen2.5 from B to B on a -sample mixture of domain knowledge and safety-alignment data, and evaluate configurations ( at three seeds) on a -item suite: what decides whether safety survives is the adapter's rank, with capacity setting how much is lost rather than whether. Under a conservative recipe, nothing below rank leaves the band around its own base, while the rank- cells at B, B and B lose – under the rule scorer and survive multiple-comparison correction; an aggressive recipe — three epochs at twenty times the learning rate — collapses every scale instead, by –. The corpus is the weaker lever: at the same rank a safety-neutral corpus with no harmful or security text still costs most of the same loss, and at rank no corpus we ran costs anything of that size. The cost is also protocol-dependent: the collapse disappears when the adaptation safety prompt is present at evaluation, so a benchmark run with that prompt can certify a model that this suite scores as collapsed. Three other results are negative and, we think, the most reusable: adaptation buys almost no knowledge, our own refusal probe mis-specifies the geometry at the top of the range, and each explanation the collapse suggests fails its own test. In practice: monitor rank and update magnitude, and keep knowledge out of the weights where retrieval can supply it.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.