Harmony in Diversity: Multi-domain Contrastive Policy Optimization for Large Reasoning Models
Abstract
Post-training via Reinforcement Learning (RL) has enabled Large Reasoning Models (LRMs) to achieve strong performance in individual domain. However, real-world applications increasingly require general-purpose reasoners rendering strong performance across diverse domains. Mixed-domain post-training aims to achieve this goal by jointly training with mixed domain data, but this often induces capability compromise and degradation among different domains. Existing methods attribute this performance degradation to harmful cross-domain interactions and propose various strategies to mitigate them, but these strategies may also impede beneficial knowledge sharing across domains and in turn fail to match or surpass single-domain performance. To address this problem, we propose Multi-domain Contrastive Policy Optimization (MCPO), which uses contrastive learning to utilize both positive and negative cross-domain interactions for knowledge sharing and competition. Specifically, we partition each rollout generated by LRMs according to its underlying reasoning structures and use reasoning segments to capture these structures. We thus formulate positive and negative pairs of reasoning segments as mutually augmented examples, which provide supportive and competing signals for knowledge sharing. Subsequently, we design complementary contrastive objectives for cross-domain knowledge sharing and intra-domain knowledge consolidation, targeting compatibility across domains and discriminability within each domain to form a harmonious reasoning space. Experimental results across a broad range of domains show that MCPO alleviates performance degradation caused by mixed-domain training and outperforms single-domain training in most cases.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.