acceptodds
Under review as a conference paper at ICLR 2027

TSRM: Two-Sided Reliability Modeling for Collaborative Reinforcement Post-Training

Abstract

Collaborative reinforcement post-training allows large language models to learn from distributed small-model evaluations without collecting their local training data. Existing approaches turn client evaluations of generated reasoning into rewards, making both the reasoning candidates and their evaluators part of the training signal. However, without reference-answer rewards, the resulting supervision is subject to uncertainty in both candidate reasoning quality and the suitability of client evaluators for different subjects. We propose TSRM, a two-sided reliability modeling framework that addresses these uncertainties by coordinating candidate selection and evaluator routing within a shared feedback loop. To guide candidate selection without correctness labels, we design MultiGranularity Candidate Confidence Filtering (MG-CCF), which captures overall confidence and local low-confidence patterns through multi-granularity aggregation of probability- and entropy-based token scores. To adapt evaluator selection to subject-dependent feedback, Stability-Aware Subject-Profile Routing (SASPR) maintains client–subject profiles through smoothed agreement estimates and variability penalties. Full-group feedback updates the profiles for subsequent evaluator selection, whereas feedback on retained candidates drives policy learning and thus shapes the next round of candidates. Experiments across six mathematical-reasoning benchmarks show that TSRM achieves the highest average accuracy under both Non-IID and identity-homogenized evaluator settings, outperforming FedGRPO-U by 2.7 and 2.2 points, respectively, without referenceanswer rewards or predefined expert assignments.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.