MOMO: Mixture of Model Optimization
Abstract
Reinforcement learning with verifiable rewards (RLVR) can train a language model to improve its solutions on a single problem. But each model refines toward its own set of behaviors, and once it stalls, further training rarely changes its approach. We propose Mixture of Model Optimization (MOMO): instead of training one model longer, we use the solutions it has found as the starting point for training other models. Each recipient keeps its own weights, so it brings different learned behaviors to the same solutions. Across mathematical constructions, circuit placement, and qubit routing with three models of roughly 30B parameters, MOMO consistently improves on single-model training, making progress after the previous model had stalled. MOMO sets a new state of the art on circuit placement, qubit routing, and the Erdos minimum overlap problem, where its gains consistently hold across seeds. Code is available at https://anonymous.4open.science/r/momo-anon-release-38B5.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.