How LoRA Remembers? A Parametric Memory Law for LLM Finetuning
Abstract
Large Language Models (LLMs) must continuously learn and update knowledge to remain effective in dynamic real-world environments. While Low-Rank Adaptation (LoRA) is widely used for such memory updates, existing studies mainly rely on qualitative downstream evaluations, leaving the quantitative capacity limits and underlying dynamics of exact parametric memory largely unexplored. To bridge this gap, we employ LoRA as a controlled memory capacity probe within the latent space to systematically quantify exact parametric memory. We introduce the LoRA Parametric Memory Law, a robust power law that characterizes incremental, latent-space parametric memory realized via LoRA, linking loss reduction to effective parameters and sequence length. At the token level, fine-grained analysis characterizes a transition from competition-dependent to probability-guaranteed recall, showing that a prediction probability of constitutes a sufficient condition for token-level recall under greedy decoding. Driven by these insights, we introduce MemFT, a threshold-guided optimization strategy that dynamically redistributes the training budget toward sub-threshold tokens. Empirical evaluations demonstrate that MemFT can enhance memory fidelity and efficiency.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.