acceptodds
Under review as a conference paper at ICLR 2027

Static Gradient Attribution Underperforms a Density-Matched Random Mask Within LoRA’s B-Matrix

Abstract

Parameter-efficient continual learning updates a restricted subset of adapter parameters to incorporate sequential tasks while limiting catastrophic forgetting. Gradient attribution is a common rule for selecting that subset, but it requires an additional backward pass over calibration data. We identify a parameter-pool confound in comparisons between initialization-time gradient masks and unrestricted random masks under standard zero-initialized LoRA (): the gradient with respect to is identically zero (), so gradient masks select strictly from the matrix whereas unrestricted random masks draw from both and . Such a comparison changes both the selection rule and the candidate parameter pool. Isolating the selection criterion by restricting both approaches to the same parameter pool, we find a negative result: a uniform random mask consistently achieves lower loss-based backward transfer than gradient attribution across seven seeds on Qwen2.5-3B. The effect is concentrated on the medical task, which contributes 80% of the aggregate; code and math move in the same direction, and legal shows no reliable difference. This advantage is consistent with a stability–plasticity trade-off under the tested task sequence, in which random masking slightly reduces single-task fitting. On the two discrete-option tasks (Medical and Legal), the mean accuracy-BWT difference is not statistically distinguishable from zero. The 95% interval extends to percentage points in the gradient mask's favour, so the pre-specified 4-point non-inferiority criterion is not met. On loss-based forgetting, Random-B consistently outperforms the density-matched gradient mask across seven seeds; on accuracy, the evidence remains inconclusive at . The attribution pass is therefore not justified by the loss-based forgetting evidence, while its accuracy consequences require a better-powered evaluation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.