acceptodds
Under review as a conference paper at ICLR 2027

SIMPLE: Verifiable Student-Improvement Feedback for Pretraining-Time Learning by Teaching

Abstract

Learning-by-Teaching (LbT) requires feedback on whether a teacher helps a student, yet raw pretraining text provides no such feedback. Neither token-level proxy gaps nor hint validity alone can provide this feedback signal, while raw student improvement can be inflated by answer leakage. We introduce Student Improvement Mode of Pseudo Learning by Teaching (SIMPLE), a framework that constructs source-grounded teaching events and derives teacher-training feedback from answer-masked likelihood gains in frozen students. SIMPLE measures teaching utility as answer-masked gains in frozen students' answer log-likelihood, and grounding and anti-hacking checks decide which hints may receive credit at all. Coverage correction assigns zero utility to events without a valid candidate, so the metric jointly accounts for valid-hint availability and the usefulness of selected hints. The teacher learns through supervised hint imitation followed by reward-filtered preference refinement that prioritizes more useful hints over valid but less useful alternatives. Across eight benchmark-derived evaluations and two small-model backbones, SIMPLE attains the highest coverage-corrected teaching utility among the evaluated methods.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.