Beyond Valid Outputs: Auditing Success Preservation in Test-Time LLM Fusion
Abstract
Combining specialist language models without further training does not guarantee preservation of their existing successes. We study this question through an empirical comparison of frozen expert groups, pairing native task accuracy with retained, lost, and additional successes. WordGame provides deterministic text-state-machine tasks and a public task-to-expert mapping. On its 768-input development panel, syntax constraints increase parseable outputs from 256 to 754, but final successes only from one to two. A separate, frozen 96-input confirmation with the same specialist bank reproduces a large weighting-policy gap: fixed task-aware weighting solves 93 inputs and dynamic weighting 46, despite all responses being parseable and uncapped. A four-policy contrast recovers 93 successful cases when the main weight is fixed at 0.75; equalizing auxiliary weights while retaining the dynamic main weight recovers 73 cases. These are effects of complete generation policies, including their fallback behavior, rather than universal evidence against dynamic fusion. With common-parent math, coding, and instruction specialists, EMFuse, CoRE, and Mean solve 5–15 code problems outside the saved specialist union yet underperform the coding specialist by 3.75–6.75 percentage points. Together, these results show that admissible outputs and additional coverage can coexist with poor preservation. The study identifies concrete policy-dependent tradeoffs without proposing a new fusion algorithm or imposing a universal method ranking.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.