Measuring Answer Accuracy and Trace Verifiability in Mathematical Reasoning
Abstract
Reasoning traces expose intermediate statements, but accepting a trace and obtaining a correct answer are distinct evaluation outcomes. We built a deterministic checker for a small model's math reasoning traces and found that checkability and correctness come apart: a trace the checker accepts is correct only about one time in three. The model was trained by outcome-only reinforcement learning to write a compact trace language whose equations and checks a program evaluates. We measured answer accuracy and checker acceptance over the same outputs, using a checkpoint panel and seed-paired runs against free-form reasoning at matched token budgets. At matched training-example exposure, the trace interface raises acceptance by about 34 points and lowers accuracy by about 16 points relative to free-form reasoning. The directions of these differences are already present at the first measured checkpoint. Across the model's own checkpoints, the accuracy–checkability trade-off we pre-registered did not confirm. Checkability and correctness are different quantities and should be measured separately.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.