SCRATCH: Efficient Math Inference for Large Reasoning Models
Abstract
Extended reasoning traces can contain repeated calculations, explanations, and abandoned branches that increase output length without contributing to the final answer. Inspired by human scratch work, we propose SCRATCH, a curriculum fine-tuning framework for concise mathematical reasoning. Scratch Trace Construction (STC) selects intact units containing computations, conditions, and justifications from native traces with verified final answers. Progressive Compression Curriculum (PCC) progressively shifts training exposure from less compressed to more compressed targets. Across three models and three mathematical benchmarks, SCRATCH reduces generated tokens by 24.87% on average, with a mean accuracy change of percentage points relative to the unadapted models. Peak savings reach 37.71% on MATH-500 with DeepSeek-1.5B. Against five efficient reasoning methods, it achieves the highest accuracy in eight of nine settings and the shortest outputs in eight. Controlled ablations show that progressive ordering improves accuracy and shortens outputs across all three mathematical benchmarks when training targets, effective batches, cumulative sample exposure, and the final consolidation stage are held fixed. Without additional task-specific training, SCRATCH reduces generated tokens on the transfer benchmarks by 34.52% on average, with a mean accuracy change of percentage points.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.