TSRM: Two-Sided Reliability Modeling for Collaborative Reinforcement Post-Training
Abstract
Collaborative reinforcement post-training allows large language models to learn from distributed small-model evaluations without collecting their local training data. Existing approaches turn client evaluations of generated reasoning into rewards, making both the reasoning candidates and their evaluators part of the training signal. However, without reference-answer rewards, the resulting supervision is subject to uncertainty in both candidate reasoning quality and the suitability of client evaluators for different subjects. We propose TSRM, a two-sided reliability modeling framework that addresses these uncertainties by coordinating candidate selection and evaluator routing within a shared feedback loop. To guide candidate selection without correctness labels, we design MultiGranularity Candidate Confidence Filtering (MG-CCF), which captures overall confidence and local low-confidence patterns through multi-granularity aggregation of probability- and entropy-based token scores. To adapt evaluator selection to subject-dependent feedback, Stability-Aware Subject-Profile Routing (SASPR) maintains client–subject profiles through smoothed agreement estimates and variability penalties. Full-group feedback updates the profiles for subsequent evaluator selection, whereas feedback on retained candidates drives policy learning and thus shapes the next round of candidates. Experiments across six mathematical-reasoning benchmarks show that TSRM achieves the highest average accuracy under both Non-IID and identity-homogenized evaluator settings, outperforming FedGRPO-U by 2.7 and 2.2 points, respectively, without referenceanswer rewards or predefined expert assignments.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.