PAIR: Pool-Adapter Inference for Routing in Training-Free LoRA Composition
Abstract
Large language models are increasingly deployed with growing libraries of task-specific low-rank adaptation (LoRA) adapters that share a backbone. A practical router must operate without task labels or examples, accommodate newly arriving adapters, and avoid scoring the entire pool for every input. We introduce PAIR (Pool-Adapter Inference for Routing), a training-free hierarchical router that treats a LoRA library as a searchable index rather than a flat list. PAIR constructs weight-only fingerprints from each adapter's last-block LoRA factors and organizes them with an online Dirichlet-process Gaussian mixture. New adapters are incorporated from their LoRA weights alone; no learned router is trained or retrained. For routing, a single adapter-free backbone forward produces the hidden state used to rank cluster representatives. PAIR then evaluates additional candidates only within selected clusters and composes the selected adapters with a calibrated compatibility score. We evaluate PAIR on 26 datasets, three 7B/8B backbones, and three backbone-specific pools of 260 adapters. On LLaMA-3.1-8B, PAIR achieves a five-family macro score of , versus for the strongest evaluated non-PAIR baseline. Holding top- uniform composition fixed, hierarchical routing alone improves the macro by points while reducing adapter-score evaluations from to a median of per input. In an isolated single-threaded CPU benchmark, median per-adapter insertion time is – ms across the tested pool sizes of – adapters, with no backbone forward.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.