acceptodds
Under review as a conference paper at ICLR 2027

WHAT SHOULD A SELF-DISTILLATION TEACHER SEE? SOLVER FEEDBACK FOR LOGICAL REASONING

Abstract

A theorem prover can check a language model’s program and explain what went wrong. Self-distillation lets a teacher copy of the model see privileged context during training while the student never sees it at test time. What should the teacher see? We test three answers on ProverQA with Qwen2.5-7B-Instruct, using Self- Distillation Policy Optimization (SDPO) after a GRPO warm start. A sibling’s verified program ties GRPO in a matched 30-step pilot, +0.67 points [−3.17, +4.50] on a 600-item holdout, and trails it by 4.60 points after 150 steps at the same mini-batch size; moving the program between prompt slots changes accuracy by +0.67 points [−0.98, +2.31] in this seed. Prover diagnostics cost 18.07 points against the program and teach the defect they complain about. An uninformative context collapses program generation: a constant string in a 150-step run, and an empty slot in a fresh matched pilot that scored 0 of 600 on each of two datasets with everything but the payload fixed. A pre-run probe of the teacher’s preference margin ranks the verified program first, so a context could be screened before a run, though the ranking depends on showing the teacher the failed attempt. But every added context scores below the plain prompt, so the probe is a relative screen, not a certificate. Every arm has one seed and one model. The pilot holdout is the clean evaluation; the 1,500-item test set was reused in development.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.