Beyond Correct Labels: Feedback, Efficiency, and Coverage
Abstract
Mathematical self-evolution requires supervision that improves answer accuracy while preserving the breadth of problems a model can solve. We propose an oracle-based mathematical co-evolution framework that constructs training targets through program execution and screens their agreement with problem statements. To audit the target–grader–policy interface, we define Active Error Exposure (AEE) as the joint probability of an incorrect target and a nonconstant task-reward group, and pair it with replay diagnostics for lost reference distinctions. The first round raises pass@1 from 36.4% to 64.6% on MATH-500 and from 17.0% to 28.5% on a fixed 500-question Olympiad subset, while retaining high pass@256 coverage. Across three recorded R2 pairs, pass@256 gains over an Agent0 reproduction are 2.4–4.2 points on MATH and 9.2–11.0 on Olympiad; native 8K/16K Olympiad checks do not remove the gap. A label-only intervention further shows a training-side result: changing only conflict targets changes delivered group-relative feedback: E conflict groups are mostly all-positive, whereas M groups are mostly all-negative. The downstream effect of that intervention remains unidentified under the available grading versions. In a native 16K Olympiad check, Oracle-based reaches 64.15% pass@32 at 786 tokens per response, versus 56.87% for Agent0 at 9,905 tokens, a descriptive 12.6-fold realized-token contrast. For Agent0, extra decoding is net negative: 276 decisions are rescued but 409 change from positive to negative from its 4K prefix to 16K.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.