acceptodds
Under review as a conference paper at ICLR 2027

LoRA's ARK: Towards an Iso-Budget Rank Scaling Law with Adaptive-Rank Kernels

Abstract

Low-Rank Adaptation and its variants dominate parameter-efficient fine-tuning, yet practitioners still have to guess one hyperparameter: the rank r. A rank too low structurally limits the update; a rank too high buys headroom only by paying proportionally more parameters and memory. Behind both horns lies one coupling: the same r sets how expressive the update can be and how many parameters it takes, so no setting of r improves one without hurting the other. Existing solutions either expand rank through fixed parameterizations that lack task-specific adaptivity, or achieve adaptivity only downward, by pruning a rank budget that must be provisioned up front. We argue that the numerical rank should instead be a property that each module learns under a fixed parameter budget. To this end, we propose Adaptive-Rank Kernels (ARK). Instead of selecting kernels by traditional Mercer validity, which only bounds the attainable rank from above, we select them by the rank they measurably induce in an update synthesized from low-rank factors. We then train the kernel's own scale as a learnable rank controller, so that different modules settle at different numerical ranks under a fixed trainable-parameter count: an iso-budget rank scaling law. On language understanding (eight GLUE tasks with BERT) and mathematical reasoning (GSM8k with Qwen2.5), ARK matches or exceeds the PEFT baselines at a matched trainable-parameter budget. It averages 81.47 on GLUE, 5.3 points above LoRA and 3.1 above the strongest baseline at the same budget, statistically indistinguishable from full fine-tuning (81.72) while training about 1% of the parameters, and it leads the matched-cost methods on GSM8k at both model scales. Measuring where the updates actually settle shows the gain comes with the mechanism we asked for: under one fixed budget the attained numerical ranks spread widely across modules and tasks instead of standing at r, so the rank of the update and the parameter count are decoupled in practice. The code will be publicly available.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.