Leveraging Instruction Tuning and Merging for Reasoning Model Adaptation
Abstract
Reasoning language models (RLMs) demonstrate impressive performance by leveraging test-time compute in the form of reasoning tokens. However, this behavior makes adapting RLMs to new domains challenging and expensive. The reason is that further training can disturb the learned behavior and degrade model performance. This makes it difficult to leverage supervised fine-tuning data with human-written solutions: although it contains high-quality annotations, it lacks reasoning tokens. In this work, we show how, despite this challenge, such data can be used efficiently for RLM adaptation. For this, we first use standard instruction tuning. Next, we leverage model merging to combine the instruction-tuned model with the original RLM, picking the merging ratio such that the resulting model's reasoning behavior on the target domain is recovered. We evaluate our method across four RLMs on coding and text summarization tasks, where it improves target-task performance by up to 11.0% while preserving reasoning behavior and limiting the out-of-distribution score degradation to on average 0.7%. Importantly, our adaptations are efficient and economical, costing less than USD $10 per model.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.