acceptodds
Under review as a conference paper at ICLR 2027

Selective Continuation Supervision and Learned Stopping

Abstract

Preserving content targets and complete-stop supervision need not preserve learned stopping. Controlled selective-repair experiments vary only which observed non-stop loss terms remain active, keeping examples, content and complete-stop losses, denominator, and optimization budget fixed. Retaining 820 high-non-stop-loss positions among 8,192 eligible positions improves flat ROUGE-L by .0694 over matched non-high-risk random support and by .0679 over conditional-token-difficulty support on 256 articles. Removing those positions produces the complementary loss of overlap. On 1,536 articles, high-loss retention remains .0571 above full-pool random support and .0007 below full support; no noninferiority margin was specified. Crossing stopping rules between the same removal endpoints reproduces a +.0675 overlap difference against a total +.0686, linking the support intervention to learned stopping. A training-selected global EOS bias recovers .0483 ROUGE-L but remains .0228 below donor stopping. Automatic source-support declines as overlap recovers. Additional controls delimit the explanation: matching initial scalar non-stop pressure does not reproduce the large repair; native removal and XSum do not establish a high-loss advantage; and the same selected supports have a smaller contrast from a native parent. The contribution is a controlled account of continuation supervision in the tested repair regime. Parameter-level mechanisms, factuality gains, and improvements over ordinary native training remain unestablished.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.