When Does In-House Adaptation Cost Safety? A Capacity-Rank-Recipe Map
Abstract
Fine-tuning an assistant on an in-house security corpus looks like the benign case: domain knowledge padded with ordinary alignment data. It is not. We fine-tune Qwen2.5 from B to B on a -sample mixture of domain knowledge and safety-alignment data, and evaluate configurations ( at three seeds) on a -item suite: what decides whether safety survives is the adapter's rank, with capacity setting how much is lost rather than whether. Under a conservative recipe, nothing below rank leaves the band around its own base, while the rank- cells at B, B and B lose – under the rule scorer and survive multiple-comparison correction; an aggressive recipe — three epochs at twenty times the learning rate — collapses every scale instead, by –. The corpus is the weaker lever: at the same rank a safety-neutral corpus with no harmful or security text still costs most of the same loss, and at rank no corpus we ran costs anything of that size. The cost is also protocol-dependent: the collapse disappears when the adaptation safety prompt is present at evaluation, so a benchmark run with that prompt can certify a model that this suite scores as collapsed. Three other results are negative and, we think, the most reusable: adaptation buys almost no knowledge, our own refusal probe mis-specifies the geometry at the top of the range, and each explanation the collapse suggests fails its own test. In practice: monitor rank and update magnitude, and keep knowledge out of the weights where retrieval can supply it.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.