Revisiting Mathematical Reasoning Verification with Deep-Research-Grounded Trajectory Distillation
Abstract
Large language models (LLMs) have shown strong mathematical reasoning capabilities, yet reliably verifying existing reasoning traces remains challenging. A reliable verifier should detect genuine errors, localize the authoritative first error step, and avoid unsupported judgments on valid reasoning. We propose Deep-Research-Grounded Trajectory Distillation (DRGTD), a framework that combines structured verification supervision with conditional localization optimization. We first construct a multi-agent teacher–reviewer trajectory-construction pipeline centered on a Deep-Research Teacher. The Teacher performs source-grounded semantic analysis and selective external mathematical checks to produce structured verification trajectories. Fail-closed trajectory validation then retains supported information and compiles it into Decision, First-Error, Process, and Unified-Deployment supervision. We next apply two-stage trajectory distillation. Phase 1 performs multi-view capability learning, while Phase 2 aligns the learned behavior with the natural deployment interface. Finally, we freeze the SFT reasoning-trace-level CLEAN/ERROR decision gate and apply frozen-gate RL only to conditional first-error localization using pairwise hard negatives and frozen-reference KL regularization. On the primary Qwen2.5-Math-7B backbone, Strict Verification Accuracy improves from 29.33% for Base to 35.33% after SFT and 37.00% after RL. BaseSFT also improves Strict Verification Accuracy and Exact FE across 10 open-weight backbones. On three representative backbones, Pairwise PG + KL achieves higher Exact FE, FE|Detected, and Strict Verification Accuracy than PPO and Bandit-OPD. Teacher-side confirmatory evaluation shows a directional but not statistically significant improvement over native verification. Overall, our results support a controlled path from Deep-Research teacher verification and validated trajectory distillation to frozen-gate conditional localization optimization for compact mathematical reasoning verifiers.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.