Where Can Adapter Composition Help? Base-Model Margins Localize LoRA Complementarity
Abstract
When is it worth composing LoRA adapters instead of using the base model alone? For multiple-choice evaluation with a pool of adapters trained on the same base model, an oracle that may fall back to the base model gains only where the base model errs; we ask which base errors a pool can fix, and whether they can be found from the base model before any adapter is run. On all 14,042 MMLU test items, the Qwen3.5-9B base model is more accurate than each of its four adapters, yet the pool holds 8.28 points of oracle headroom, and the share of base errors that some adapter fixes falls from in the lowest base-margin quartile to in the highest. The same decline appears for independently trained Llama-3-8B adapters and for a new pool with a different domain mix, and on five tested workloads, where adapters change base answers, observable without labels, tracks where their rescues lie. A threshold frozen from 400 labels before the new pool was trained keeps 91.7% and 94.0% of its rescues on MMLU and MMLU-CF while sending about a third of the items to the adapters. Refitted post hoc after eight numeric calibration labels are mapped to letters, the threshold keeps 90.0% and 92.1% of these rescues with fewer passes. These rescues are the opportunity left to any selector; an exact decomposition of router improvement into this ceiling, missed gains and damage, together with a margin ceiling and a perturbation condition, serves as the accounting.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.