Taming Attention to Compose Soft Prompt Experts
Abstract
Large language models can be adapted into specialized experts through parameter or soft-prompt tuning. Combining independently trained soft-prompt experts without joint retraining is challenging because direct joint readout can misbalance both the expert group's total influence and the allocation among experts. We propose TAME(Two-level Attention Mixing of Experts), a training-free, input-conditioned composition method. TAME uses two-level attention to separately control the expert group's total readout and the within-group contribution of each expert while keeping all experts independently available. On GLUE, using the same set of soft-prompt experts, TAME achieves 94.4% mean capability retention, improves direct joint reading by 11.4 percentage points, and outperforms the evaluated context-space composition baselines. In a unified comparison with parameter merging, TAME achieves 1.90 percentage points higher GLUE Macro accuracy despite a lower single-expert reference, matches or exceeds capability retention on mathematics, code, and instruction following, and exhibits complementary capability-retention patterns with parameter merging. Overall, TAME provides a competitive training-free route for composing soft-prompt experts.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.