Don't Train My Model: Domain-Conditioned Training Resistance in Open-Weight LLMs
Abstract
Release-time evaluations of open-weight language models characterize their predictions before adaptation, but not how they respond to subsequent fine-tuning. We study whether a checkpoint can preserve these predictions while becoming harder to fine-tune on a specified domain. We introduce DormantBomb, a checkpoint transformation that rewrites a small number of existing SwiGLU coordinates by pairing target-sensitive activations with zero down-projection columns. The resulting dormant paths have zero forward contribution at release while allowing large initial gradients during downstream adaptation. In the primary Qwen3-4B experiments, DormantBomb largely preserves release-time predictions while selectively slowing target-domain fitting, with little change under separate controldomain training. This separation holds under LoRA, direct down-projection training, and full fine-tuning. In unsafe fine-tuning experiments, differences in trainability are also accompanied by differences in generated behavior.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.