acceptodds
Under review as a conference paper at ICLR 2027

Compositional RL: Evaluating and Training LLMs for Collaborative Reasoning

Abstract

Language models are often evaluated alone but increasingly deployed together, routed between, made to debate, ensembled, or collaboratively trained. For models inside systems, what matters most is not how well it performs on its own, but how much the system gains from its presence. We call this a model’s compositional strength, formalize it as the model’s Shapley value in a cooperative game over coalitions of a model pool, and estimate it with leave-one-out marginals costing n+1 system runs rather than 2n. Evaluating both quantities across five model pools, twelve collaboration methods, and twelve datasets, we find them largely uncorrelated: in 12 of 15 (pool, method-family) settings individual strength bears no significant relationship to compositional strength, and the coupling is weakest among general-purpose models that practitioners are most likely to deploy. Individual benchmarks and optimizing for them thus mean little for models’ values in collaborative systems. We then propose compositional reinforcement learning (C-RL), which places a policy inside a live multi-model collaborative system and rewards it with that system’s outcome rather than its own correctness. The leave-one-out baseline in Shapley values cancels under a group policy gradient, so this costs no more than ordinary RL and never requires running a coalition without the policy. Across three collaboration protocols and eight tasks, C-RL produces the strongest collaborator in 21 of 24 settings, is the only method that never leaves the system worse off than the untrained policy, and carries over to unseen datasets, unseen collaboration protocols, and pools of frontier proprietary models.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.