The Behavioral Rate of Fine-Tuning: Compression as a Measure of Useful Information
Abstract
Low-Rank Adaptation (LoRA) adapts a frozen language model through a compact learned update, making it a natural setting for asking how much information fine-tuning actually needs to write. Yet adapter rank, parameter count, or numerical reconstruction error do not directly tell us how many bits are required to preserve the behavior acquired during training. We therefore define an operational behavioral rate as the exact serialized size of a compressed adapter. It is the smallest number of bits that, once reloaded into the frozen base model, retains a fixed fraction of the adapter's held-out fine-tuning gain. Across a broad set of LoRA fine-tuning experiments, we find that this rate increases with distinct training content at fixed optimization budgets, but is not determined by source description length, fine-tuning gain, or ordinary weight-space reconstruction error. Most importantly, the same datasets exhibit substantially different, and even reordered, rates across frozen base models, showing that fine-tuning compressibility is a property of the dataset–receiver pair rather than of the dataset alone. Receiver-relative correction spectra are our strongest predictor of behavior-preserving rate across models and tasks, while rank and layer controls reveal an additional dependence on the adapter carrier. Finally, in controlled experiments with deliberately corrupted reasoning supervision, aggressive compression can recover useful task performance from otherwise harmful updates, beyond what can be explained by simple shrinkage, with substantially stronger recovery when more adaptation rank is available during training. Together, these results suggest that the information cost of fine-tuning is not an intrinsic property of the data or of parameter displacement, but a receiver- and behavior-relative property of how the required corrections are represented in the learned update.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.