acceptodds
Under review as a conference paper at ICLR 2027

Closing the Self-Teaching Loop for Open-Ended Generation: Sibling-Conditioned Distillation Meets Self-Verification

Abstract

Self-teaching improves language models by learning from their own generations and assessments, but extending this loop to open-ended generation remains challenging because response quality is multidimensional and task-native verification is generally unavailable. Rubrics provide a natural specification in this setting, but do not by themselves close the learning loop: how should a fixed rubric be grounded in the model's current responses, and how should stronger completed responses be distinguished from weaker ones? We introduce , which combines local and global routes to address these challenges. The local route uses same-prompt siblings to ground rubric guidance in current-policy realizations, while the global route reuses rubric-conditioned forward pass to derive continuous response scores for group-relative optimization. Across Qwen3-4B and Qwen3-8B on medical and science tasks, S²PO improves over the base model by 3.09–5.26 points and outperforms both RGSD and 32B-judge GRPO in all four in-domain settings. These gains persist under out-of-distribution evaluation, with improvements of 2.46–5.14 points over the base model across both model scales.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.