Peer-Critique On-Policy Distillation from Heterogeneous Models for LLM Reasoning
Abstract
Improving an already-capable reasoning model requires repairing its remaining failures. Verifiable rewards can reinforce successful exploration but do not explain how a particular failed trajectory should change; a critique can provide that problem-specific guidance. The model may also fail to diagnose its own error; one remedy is a much larger critic, which adds cost and presumes access to a stronger model. We instead replace the single critic with a panel of same-size peers from different families, whose critiques repair complementary failures: three 3-4B peers repair 41.8-66.0% of failures, exceeding either of two 27-32B critics on every target. **PECO** (**Pe**-**C**ritique **O**n-policy Distillation) uses the target's frozen base, conditioned on a failed attempt and a peer critique that enabled a verified repair, as the teacher for on-policy distillation. Already-solved problems enter without critiques, anchoring training to the initial policy to limit forgetting. The construction requires no reference rationale or model larger than the panel; only the verifier accesses the gold final answer. PECO achieves relative gains of 6.4% over the base models and 6.9% over on-policy self-distillation, beats the verifiable-reward baseline on three targets, and outperforms distillation from much larger teachers on all four. PECO-efficient prunes the critique pool, retaining 93-102% of PECO's four-benchmark accuracy with 36-44% fewer data-building TFLOPs, and uses 2.1 fewer generation calls than GRPO.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.