Mixture of Outputs: Dynamic Decoding-Time Expertise Integration Between Large Language Models
Abstract
Adapting large language models (LLMs) to new capabilities typically requires additional post-training, which can be costly in both supervision and computation. Existing approaches reduce this cost either through parameter-efficient fine-tuning or decoding-time adaptation, but the former still requires task-specific training of the target model, while the latter often relies on fixed combinations of auxiliary predictions or specially matched proxyxen when the desired capability is already available in another specialized models. Meanwhile, specialized capabilities can often be acquired more cheaply in smaller models. This motivates a question: can a smaller specialized model serve as a low-cost capability donor for a larger model? Therefore we study whether such capabilities can instead be transferred at decoding time, without updating the pretrained models themselves. To this end, we propose Mixture of Output (MoO), a lightweight framework that dynamically transfers capabilities from one or more frozen donor models to a frozen receiver. At each decoding step, MoO uses the receiver’s hidden state to determine how much each donor should contribute to the next-token prediction. Only a lightweight module is trained to predict the contribution of donors, while the receiver and donor models remain frozen. Our evaluation shows that MoO improves the receiver by up to 14.27 percentage points on average across nine benchmarks. For alignment transfer, MoO reduces ASR from 0.44 to 0.02 and from 1.00 to 0.17 on the receiver models. Combining multiple reasoning donors provides further gains of 2.01 and 1.92 points on the two tested receivers. Overall, these results demonstrate that MoO can effectively transfer complementary capabilities from specialized donor models to a frozen receiver across various tasks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.