Does Post-Training Tax Exact Arithmetic? It Depends on the Route
Abstract
Does post-training reduce accuracy on exact arithmetic, or can the evaluation interface change that conclusion? We study 16 released checkpoint pairs on six integer-arithmetic tasks, including seven documented post-training comparisons from three pretrained bases. We first compare checkpoints using the same completion input, without a chat template. We then compare completion and chat inputs for the same post-trained checkpoint. The common comparison between base completion and post-trained chat changes both checkpoint and interface. In a pre-registered replication that constrains outputs to integers, the common and fixed-completion comparisons give opposite conclusions in 24 of 96 pair–task combinations: one indicates significantly higher accuracy and the other significantly lower accuracy after Holm correction. In tokenizer-compatible pairs, the size of the checkpoint effect depends on the input interface. Two selected problem subsets show reversals under both constrained decoding and free generation with first-integer scoring. This does not establish independence from generation and scoring conditions. Some deficits are no longer detected after changing the prompt, whereas others persist. For OLMo-3, we ask subtraction questions without a chat template and allow free generation, scoring the first integer. On 14-digit problems, accuracy is 52.4% for the base checkpoint and 1.4% for the final checkpoint. Among answers that are correct at the base checkpoint but wrong at the final checkpoint, most first become wrong at the supervised fine-tuning stage. For Llama-3.1/Tulu, gains measured with the same completion input become negative estimates under the question form. A secondary transplant of base-checkpoint states corrects some selected expression-form errors but none of ten tested 14-digit subtractions. These findings concern accuracy under specific evaluation conditions. Comparisons of post-training should distinguish checkpoint, interface, and answer-extraction conditions before attributing score differences to changes in arithmetic ability.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.