acceptodds
Under review as a conference paper at ICLR 2027

Policy-Merge: Policy Prediction for Multi-Domain RL and Distillation

Abstract

Language models are post-trained with Reinforcement Learning (RL) across many domains, such as math, code, and tool use. The resulting model depends on how data and compute are allocated across these domains, whether by training on a data mixture or by multi-teacher distillation. Evaluating an allocation, however, requires a full training run. Existing mixture selection methods reduce this cost by predicting benchmark performance or using proxy models for mixture optimization, but do not directly predict the policy obtained by training on a mixture. We introduce a method that approximates the policy of a model trained with RL on a given data mixture by composing independently post-trained domain experts at decoding time. We compare this approximated policy with the model trained with RL on the mixture, and show that they perform similarly across benchmarks and are distributionally close. This allows for predicting the capabilities of a specified mixture. Our method enables efficient data mixture selection, as candidate mixtures can be evaluated and ranked in policy space before committing compute to post-training. Finally, we show that a student distilled on-policy from several experts at once to minimize a weighted sum of per-token KL divergences is equivalent to using our approximated policy as a teacher. One can thus skip the multi-teacher distillation and decode the approximated policy directly.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.