acceptodds
Under review as a conference paper at ICLR 2027

MOMO: Mixture of Model Optimization

Abstract

Reinforcement learning with verifiable rewards (RLVR) can train a language model to improve its solutions on a single problem. But each model refines toward its own set of behaviors, and once it stalls, further training rarely changes its approach. We propose Mixture of Model Optimization (MOMO): instead of training one model longer, we use the solutions it has found as the starting point for training other models. Each recipient keeps its own weights, so it brings different learned behaviors to the same solutions. Across mathematical constructions, circuit placement, and qubit routing with three models of roughly 30B parameters, MOMO consistently improves on single-model training, making progress after the previous model had stalled. MOMO sets a new state of the art on circuit placement, qubit routing, and the Erdos minimum overlap problem, where its gains consistently hold across seeds. Code is available at https://anonymous.4open.science/r/momo-anon-release-38B5.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.