Perfect Labels for Free? What Exact Answers and Exact Traces Teach a Small Language Model
Abstract
Exact verifiers (primality tests, SAT solvers, unit tests) label data without noise and supply the rewards of verifiable-reward post-training. Because every label is exact, a difference between two training regimes on the same inputs is a difference in the supervision format. We fine-tune Qwen2.5 models (0.5B to 7B; main size 1.5B, three seeds for the core answer and trace cells) with LoRA on exactly labeled decision tasks and compare supervision by the answer alone with a deterministic procedure trace followed by the answer. Exact answers teach shortcuts: on 4-digit primality the 1.5B answers (three seeds) follow "the last digit is 1, 3, 7 or 9" on 99.2-100.0% of test items, and on divisibility by 7 or 13 answer-only training stays at chance or constant at every size and configuration we ran (n<=540, including a logged sweep of four configurations, eight runs, none collapsing), while the same recipe learns divisibility by 3. Exact answers remove computation from the answers: at 3B and 7B the zero-shot model already divides by 7; answer-only fine-tuning drops it to chance in both seeds, and under a chain-of-thought prompt the answer-only model still answers at once at every size tested (at least 97.9% of items in three tokens or fewer) and stays at chance, while fine-tuning on its own verifier-filtered traces keeps the division (two seeds at each size); whether the capability survives is untested. Exact traces teach steps within a task: accuracy rises with the number of traces and with how often each table entry was supervised, a count-based prediction across tasks failed because steps differ in difficulty, and only the input's own trace, written before the answer, helps. A second model family (SmolLM2-1.7B-Instruct) shows the same direction on the three tasks run, but its timestamped prediction (S11) failed its 20-point threshold in one seed. Predictions were logged in timestamped commits before their cells ran; four of twelve failed (the cross-task dose prediction S9, the length prediction S5, the self-generated-trace prediction S10 and the second-family prediction S11), and all are reported.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.