LoopedLoRA: Low-Rank Adaptation with Recurrent Subspace Depth
Abstract
Low-Rank Adaptation (LoRA) is a widely used parameter-efficient fine-tuning method. A common way to improve LoRA adaptation is to increase the rank, which also increases the number of trainable parameters; despite many variants, most adapters retain a single-pass mapping through the low-rank bottleneck. We propose LoopedLoRA, which introduces weight-shared recurrent depth within this subspace. LoopedLoRA recurrently processes an -dimensional latent state using a fusion-friendly RMSNorm–SiLU operator, adding only more trainable parameters than LoRA at on Qwen3-4B-Base. To allow the contribution of recurrent processing to vary across tokens, a learned latent blender adaptively combines the direct and recurrent paths, and LoopedLoRA-Light reuses its score to bypass recurrence at inference. Across four Qwen3 and Llama-3 models, LoopedLoRA improves GSM8K by up to 12.96 percentage points over LoRA and outperforms rank-doubled LoRA () and competitive baselines on mathematical reasoning, while remaining competitive on code generation and commonsense reasoning. With fused Triton kernels and CUDA graph replay, full LoopedLoRA maintains decode latency comparable to LoRA, while LoopedLoRA-Light achieves up to faster decoding.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.