Global Factorized Adaptation: Efficient Parameter Sharing for Fine-Tuning Language Models
Abstract
Low-rank adaptation splits the trainable budget across modules and constrains the rank of every module update, and existing sharing methods typically keep this module-wise structure or rely on fixed expansion maps. We study Global Factorized Adaptation (GFA), the simplest configuration of the SuperLoRA design space (one group, a matrix reshape, and no projection), which instead imposes low-rank structure on the model-wide update: GFA arranges the coordinates of all target updates into a near-square matrix, learns a single low-rank factorization of it, and reconstructs each module update by slicing and reshaping. This layout minimizes the parameter count for a fixed global rank and makes the budget about times finer than that of uniform rank-one LoRA over equally shaped modules. We give an arithmetic layout condition, independent of the module order and satisfied by our layouts, under which the module updates are generically full rank even at global rank one, and we show that on both evaluated models, GFA and LoRA are far from containing each other at matched budgets. Against six baselines on two language models, six tasks, and five matched budgets, GFA improves over LoRA by percentage points on average, with the largest gains at small budgets. With to fewer parameters than rank-one LoRA, it outperforms partial-module LoRA of the same size by to percentage points on both models, and its training cost stays close to that of LoRA.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.