acceptodds
Under review as a conference paper at ICLR 2027

From Coverage to Reliable Success: Resampled-Evidence Policy Optimization for Test-Time Reinforcement Learning

Abstract

We study how label-free reinforcement learning can improve sampled accuracy () while retaining broad solution coverage () and repeated correctness. The first two metrics are our primary outcomes; , the probability that four independent responses to the same question are all correct, provides a supporting reliability assessment. A central challenge is that finite verification evidence can support different responses unequally, even after a supervision target has been selected. We propose Resampled-Evidence Policy Optimization (REPO), which retains coverage-oriented training while adapting each response's contribution to its evidence support. The student and teacher maintain separate parameter states; the teacher supplies judgments and is updated only by an exponential moving average (EMA) of the student parameters. By resampling cached verification judgments and reconstructing supervision on fixed student responses, REPO attenuates less reproducible decisions without additional verification decoding. For each response, agreement and conflict between replayed and original advantage signs determine a bounded update weight. This weight preserves the original direction while reducing the influence of decisions sensitive to the sampled evidence. Across five external benchmarks, REPO improves macro-averaged over Co-rewarding by 2.52 and 1.60 percentage points at 4B and 8B, with gains of 2.12 and 2.18 points. Macro-averaged remains close to Base (+0.18/+1.06 points), despite larger gaps to Co-rewarding (+11.50/+12.00 points). These results indicate aggregate coverage retention, with losses on some individual benchmarks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.