Co-RL: Unsupervised Reasoning Emerges from Diverse Cohort in Multi-agent RL
Abstract
Reinforcement Learning (RL) has emerged as a powerful approach for improving reasoning in language and vision-language models, yet its strongest successes still depend heavily on ground-truth supervision (e.g., verifiable reward). Such annotations are costly to obtain and become increasingly scarce as reasoning capabilities advance beyond what humans can reliably evaluate. Self-rewarding RL reduces this dependence by enabling models to derive reward signals from their own completions. However, training solely on self-generated feedback can reinforce existing biases, reduce response diversity, and ultimately lead to homogenized responses and training collapse. In this work, we show that unsupervised reasoning can emerge through cooperative multi-agent training. We introduce Co-RL, a framework in which multiple decoupled models, sharing no parameters, are simultaneously optimized through RL using rewards derived from their peers. We further show that increasing cohort diversity, through heterogeneous model families, sizes, and rephrased training samples, reduces the correlated errors that drive self-reinforcing feedback loops. This diversity consistently improves reasoning performance and mitigates training collapse. Without any ground-truth labels, Co-RL improves the base models by 3.0–7.2% on average across seven benchmarks and by 1.6–4.4% across four multimodal benchmarks. Compared with prior state-of-the-art methods, it also outperforms the multi-agent RL method CoMAS by 4.0% and the self-rewarding vision-language RL method MM-UPT by 2.0%.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.