GALA: Reward-Guided Distillation Across Tokenizers for Arithmetic Reasoning
Abstract
Distilling a language model for verifiable reasoning requires reconciling two forms of supervision: a teacher's token distributions and a verifier's assessment of complete answers. When teacher and student use different tokenizers, their distributions must first be aligned; even after alignment, imitation and reward-based learning can prescribe conflicting updates. We present GALA, a training framework that combines cross-tokenizer distribution matching with group-relative verifier feedback. It constrains teacher gradients in the student's aligned logit space through conflict projection, a learned gate, and a verifier-relative norm cap, while retaining reference-solution supervision. We evaluate completed training configurations on Countdown with a Qwen3-4B teacher and Gemma-3-1B and Llama-3.2-1B students. Fixed forward achieves 44.87% and 46.12% raw accuracy in the seed-42 direction comparison; its small lead over outcome conditioning does not establish a reliable direction advantage. Supplemental component ablations report 41.00% and 41.80% strict accuracy for full GALA, exceeding unmodified KD on the same eligibility support by 5.00 and 4.90 percentage points. Fixing the gate to one while retaining projection and the norm cap yields 40.90% and 41.60%, indicating a much smaller observed contribution from learned gating. These results support gradient control over directly adding the aligned teacher gradient in the evaluated configurations, while distinguishing arithmetic correctness from successful completion.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.