LoRA Memorization Scaling Laws
Abstract
LoRA adapters have become the canonical method for injecting new knowledge and behavior into a pretrained model, yet their physical limits are not yet fully understood. This work measures the capacity of adapters to memorize and generalize by training on synthetic datasets of both random facts and hidden modular rules. On the DataDecide OLMo checkpoints between 20M and 1B parameters, we observe that an adapter stores up to 2.2 bits per projection parameter, with an efficiency that falls with rank and rises with base model quality. Our memorization scaling laws predict the target held-out 1B x rank within 11% errors. While capacity alone does not decide whether an adapter can generalize to modular rules, it requires a sufficient number of directions to represent facts and can grok faster on a pretrained base checkpoint of higher quality.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.