Queryable LoRA: Instruction-Regularized Routing Over Shared Low-Rank Update Atoms
Abstract
Low-rank adaptation (LoRA) applies the same learned update to every input, even when an adapter must perform multiple tasks. Queryable LoRA is a parameter-efficient fine-tuning (PEFT) approach that adapts its low-rank update to each input by retrieving a sparse mixture of rank-space atoms from a shared memory. A block uses its current representation and a summary of earlier blocks to select the atoms. In the Instruction-Queryable LoRA (IQ) approach, an instruction embedding also guides the query and defines a prior over atoms. We show that the router solves a KL-regularized retrieval problem. We also demonstrate that its updates stay norm-bounded across routing switches and that the blockwise gradients factorize exactly. To test transfer, we trained one adapter on a mixture of source tasks and chose its checkpoint on validation data. We then evaluated it on held-out tasks with unseen instructions. Across model sizes, Queryable LoRA matched or exceeded LoRA in most paired runs. In the main comparison, both variants exceeded every baseline in average accuracy at a training cost close to LoRA's.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.