acceptodds
Under review as a conference paper at ICLR 2027

Learning to Verify by Doing: A Self-Renewing Policy–Verifier Loop for Sustained Self-Improvement

Abstract

Verifier-guided self-training can stall as policy improvement erodes the selection advantage of existing evaluative guidance. We show that execution feedback can renew this guidance by teaching a verifier about action outcomes and applicability conditions. We introduce Knowing–Doing Co-Evolution (KDCE), a self-renewing policy–verifier loop that learns outcome prediction and candidate ranking from budgeted executions and controlled local action comparisons, then uses uncertainty-aware selection for weighted policy self-training. The updated policy supplies experience for the next round, closing the loop. Across 8 checkpoints from 3 model families and 11 code, mathematical-reasoning, and interactive evaluations, KDCE improves family- and domain-balanced standalone performance by 0.7 percentage points over prespecified domain-specific references after 8 rounds. Matched-evidence and directional interventions support both feedback paths. Gains are heterogeneous: improved verification does not always yield policy gains or compute savings.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.