SIFT-LoRA: Choose What a Layer-Sparse Adapter Inherits
Abstract
A trained LoRA adapter should be reusable across deployment storage budgets chosen after task adaptation. We introduce SIFT-LoRA, which compiles a same-task Full-LoRA parent into a layer-sparse adapter under an exact serialized-byte cap without further training or changing the retained weights. A layer’s value can depend on which other updates are retained, so fixed importance scores may miss its contribution to the sparse child. Our selector, Conditional SIFT-KL, evaluates how much each candidate layer reduces parent–child predictive KL when added to the current child. It chooses the feasible addition with the largest positive reduction per added serialized byte, then recomputes these reductions for the remaining candidates. Across three models and three tasks at 12.5% and 25% of the full adapter’s serialized size, Conditional SIFT-KL lowers predictive KL against each of the fixed-score and contiguous-layer baselines in 16 of 18 conditions. It also achieves the highest task utility among the three selectors in 14 of 18 conditions, including ties. On Qwen3.5-2B/HellaSwag at the 12.5% cap, it improves mean accuracy over fixed scoring by 3.0 percentage points across three independently trained parents at nearly identical adapter sizes. In a separate seven-method BoolQ comparison at the 12.5% cap under unmerged bf16 H100 inference, SIFT-LoRA lies on the observed utility–latency Pareto frontier for all three models. Broader deployment experiments cover five tasks and both byte budgets. These results support choosing layers by what they add to the retained set and separating task learning from deployment capacity.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.