The Convention Is the Grader's: How Model Proof Graders Price Silently Repaired Errors
Abstract
The grading rubric bands a small repairable error at 5–6 and says nothing about a false line the proof itself corrects; what it costs is the grader's judgement. We measure that price and propose no grader. Into 24 IMO 2026 proofs, model-judged closed, not human-verified, we plant three single-hunk edits for 96 texts — fatal, harmless (nothing later uses it), or a deleted argument. Falsity is machine-checked; class is a recorded label (8 of 72 disputed). A panel of 5 graders scores the texts, and every decision rule is read under two conventions for the repaired line. Graders notice harmless slips and price them differently: on the 15 harmless texts where both pair graders quote the planted slip, one awards full marks on 2/15 and the other on 14/15, reversing which grader looks most accurate on the strict mark; a retest keeps the reversal. The completeness flag sits between score thresholds and differs from them only within grade 6. An obligation audit whose prompt says not to repair, read as a comparator, orders the graders the same way (12/24 against 24/24 harmless mutants accepted under its rule); its held-out gap to the strict mark is exploratory and unresolved (p ≥ 0.18; its gain three announced deletions), and a pre-registered replication confirms only its gap to the pass mark. On human-majority-labelled competition proofs, the strict mark and flag err mostly by false rejection. With the audit's no-repair sentences in the grade prompt, the strict mark moves to 6/32 and 13/32 (exploratory).
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.