acceptodds
Under review as a conference paper at ICLR 2027

Measuring Answer Accuracy and Trace Verifiability in Mathematical Reasoning

Abstract

Reasoning traces expose intermediate statements, but accepting a trace and obtaining a correct answer are distinct evaluation outcomes. We built a deterministic checker for a small model's math reasoning traces and found that checkability and correctness come apart: a trace the checker accepts is correct only about one time in three. The model was trained by outcome-only reinforcement learning to write a compact trace language whose equations and checks a program evaluates. We measured answer accuracy and checker acceptance over the same outputs, using a checkpoint panel and seed-paired runs against free-form reasoning at matched token budgets. At matched training-example exposure, the trace interface raises acceptance by about 34 points and lowers accuracy by about 16 points relative to free-form reasoning. The directions of these differences are already present at the first measured checkpoint. Across the model's own checkpoints, the accuracy–checkability trade-off we pre-registered did not confirm. Checkability and correctness are different quantities and should be measured separately.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.