A Two-Resource View of Adapter-Library Arbitration
Abstract
Growing adapter libraries shift the central challenge of parameter-efficient continual learning from forgetting to arbitration: given an input, which independently trained adapter should answer? Across frozen-language adapter libraries with identical adapters under every arbitration rule, we identify an empirical two-resource pattern. Fitted routers depend on sufficient cross-task meta-training data relative to routing difficulty: holding adapter training fixed, reducing router labels from 1,000 to 100 lowers accuracy from 86.8% to 53.7%; jointly starving adapter training lowers it to 35.4%, while the unfitted competition reaches 83.9% in that cell. Unfitted abstention-based competitions avoid this dependence but pay a different cost: as the library grows, max-over-K arbitration amplifies the upper tail of imperfectly comparable wrong-expert scores. This leads to Reference-Anchored Slack (RAS) training for frozen-language libraries: a small shared, domain-adjacent reference trained as slack inside each adapter's existing head, producing more comparable native confidence scores without K-way fitting or refitting. The reference alone adds +6.3 ± 3.9 points to the confidence competition at K=75 on CLINC150, and its value tracks reference proximity—exchanging an adjacent reference for a distant one costs up to 11.5 points. Within the reference-trained substrate the competition's K-cost stays shallow from 125M to 32B, whereas matched no-reference controls incur larger K-costs at every tested scale, so scale alone does not ensure cross-expert comparability. For generation, the same training principle supplies the shared out-of-scope coordinate that first-token margin arbitration requires. Our results give a regime map for choosing adapter-library arbitration.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.