acceptodds
Under review as a conference paper at ICLR 2027

How Much Evidence Should a Self-Correction Carry? Adaptive Dirichlet Evidence for Self-Distillation

Abstract

Language-model agents can generate, execute, and revise their own solutions, creating supervision for self-improvement. A correction’s apparent quality, however, need not reveal how concentrated its external support is. We introduce Effective-Evidence Self-Distillation (EESD), which separates relative transition support from an effective pseudo-count mass derived from normalized execution relevance. A Dirichlet posterior converts these quantities into an uncertainty-penalized correction weight for KL-anchored self-distillation. Under a symmetric prior, changing mass preserves category ordering, while effective mass cannot increase the supervised coefficient relative to matched fixed-mass weighting. Across four model–domain history sweeps, increasing visible observations from one to eight reduces EED future-outcome NLL by 55.0–59.3%. At eight observations, EED also achieves lower NLL than fixed mass in all four comparisons. In a matched DeepSeek/RunBugRun study, the predictors agree in argmax on all 3,000 examples, with the largest NLL gain in the most concentrated relevance bin. The complete learning update achieves its largest observed downstream gain on DeepSeek/CodeARC, raising all-tests Pass@1 from 15.0% to 20.4%. These results connect an explicit evidence decomposition to probability estimation and correction learning.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.