acceptodds
Under review as a conference paper at ICLR 2027

RaSTA: Random Subspace Tuning Adaptation

Abstract

Parameter-efficient fine-tuning (PEFT) methods such as LoRA specialize a pretrained model by learning a small additive update to its frozen weights. However, LoRA's trainable parameter count is tied to the shape of each adapted matrix, costing at least the sum of its input and output dimensions even at rank one. We propose Random Subspace Tuning Adaptation (RaSTA), which adapts the model within frozen random subspaces shared across layers and learns only how their directions interact, allowing trainable parameter budgets well below rank-one LoRA without storing the frozen directions in the adapter checkpoint. Its two variants trade off the expressivity of the interactions against the number of directions reached under a matched budget. RaSTA-DM learns a dense mixing matrix between the bases, while RaSTA-HS fixes the mixing as a Walsh-Hadamard transform and learns one scaling vector per basis. To our knowledge, this is the first use of the Walsh-Hadamard transform for this purpose. The latter is somewhat more costly but performs better and degrades more gracefully as the budget shrinks. Across mathematical reasoning, commonsense reasoning, and code generation with 7-8B models, as well as natural-language understanding with RoBERTa-large, both variants remain competitive with rank-one LoRA, and sometimes with higher-rank LoRA, at budgets ranging from less than half to about one tenth of its parameter count. Both variants also train faster and use less peak memory than rank-one LoRA, except for RaSTA-HS at its largest considered budget.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.