acceptodds
Under review as a conference paper at ICLR 2027

DESA: Delivered-Energy Spectral Allocation for Merging Reinforced Experts

Abstract

Model merging provides a practical way to combine multiple experts into a single model without retraining or increasing inference cost. However, existing merging methods are largely developed for supervised fine-tuning and do not account for the distinctive structure of reinforcement learning with verifiable rewards (RLVR) updates. We show that RLVR task vectors exhibit two properties that are especially important for merging: (i) apparent sparsity under low-precision representation, and (ii) substantial heterogeneity in update energy across experts. These properties make sparsification-based heuristics unreliable and cause existing spectral methods to impose an uneven, supply-determined allocation of delivered energy, which strongly affects axis-wise recovery after merging. Motivated by this observation, we propose DESA (Delivered-Energy Spectral Allocation), a training-free spectral merging method that explicitly reallocates expert-level delivered energy while preserving a shared spectral frame. DESA introduces controlled budget equalization across experts together with retention floors that prevent under-allocation to experts whose capabilities are sensitive to reduced budget. Across both language and vision-language RLVR expert pools, DESA consistently outperforms strong baselines in average accuracy and recovery. Our results identify explicit delivered-energy allocation as a key ingredient for merging RL-trained experts.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.