Hard to Remove, Cheap to Replace: Compact, Transferable Surrogates for Sensitive LLM Layers
Abstract
Removing whole layers is a direct way to compress a large language model, but the damage it causes varies widely across layers and tasks, and removal alone does not show whether that damage can be recovered. We study layer sensitivity through recovery: how much of a removed layer’s contribution a compact learned replacement can restore, and whether that replacement transfers to tasks it was not fitted on. Using Llama 3.1 8B and Qwen3-8B, with the rest of each model frozen, we apply three interventions to complete layers during free text generation: Skip removes a layer, Replace substitutes a small surrogate fitted to the layer’s residual output, and Transfer reuses a surrogate fitted on a different dataset. We extend these interventions to pairs of layers, compare linear and nonlinear surrogates, and contrast trained surrogates with calibrated random maps. A linear surrogate of rank 64 with about 1/413 of a Llama layer’s parameters improves over Skip in all 12 sensitive Llama settings, recovering up to 93.2% of the lost performance. Surrogates also transfer across datasets and task categories: in Llama, surrogates fitted in one task can transfer to other tasks, still with significant improvement over Skip. Jointly fitting two surrogates for two layers improves the replacement over independent surrogate composition. And we prove the effectiveness of the surrogates training by showing it outperforms random maps in all tasks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.