Gradient-guided Fine-tuning
Abstract
Supervised fine-tuning (SFT) trains language models to reproduce reference responses token by token, although different reasoning trajectories can support the same correct answer. We introduce Gradient-guided Fine-Tuning (GFT), a method for learning from reference solutions through task-relevant activation alignment. GFT generates a student response and replays both reference and student reasoning through the same model. Gradients of the reference-answer negative log-likelihood identify answer-sensitive directions, while a directional activation-alignment objective guides parameter updates without directly optimizing full-response cross-entropy. We combine these activation-alignment losses across all decoder layers using normalized weights, without requiring model-specific layer selection. To mitigate late-training generation collapse, we incorporate KL regularization on student-generated contexts, preserving useful generation behavior while permitting representation-level adaptation. We evaluate GFT on three mathematical reasoning benchmarks, GSM8K, MATH Intermediate Algebra, and MATH Precalculus, across three model sizes: Qwen3-0.6B, Qwen3-1.7B, and Qwen3-4B. The best GFT configuration exceeds SFT by 0.8–3.6 percentage points on GSM8K and by 7–33 points on the two MATH subjects. These findings support gradient-guided activation alignment as a promising alternative to direct token imitation for adapting reasoning models. The code is available at https://anonymous.4open.science/r/GFT/.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.