acceptodds
Under review as a conference paper at ICLR 2027

Drift-Constrained Optimization: Only Direction Matters in Fine-Tuning Instruct Models

Abstract

Fine-tuning instruct models often improves target performance while inducing behavioral drift that can degrade existing capabilities. Rather than treating drift as an uncontrolled consequence, we impose a drift budget and ask how to maximize target performance within it. Locally, behavioral drift defines a geometry around the reference model, where the budget constrains distance, leaving direction as the remaining degree of freedom. Directional efficiency thus reformulates fine-tuning as a direction-selection problem, predicting that changes in accessible directions can qualitatively alter outcomes. We test this prediction in a stringent QA-only setting, where strong instruct models are trained only on final answers yet must generate multi-step reasoning at inference. A coarse layer-selective probe reverses QA-only fine-tuning failure, with neighboring configurations improving target performance while preserving reasoning and general capabilities. Across Qwen3 models, these directions substantially improve scientific reasoning and multilingual translation. Over 100+ languages, the resulting models match or outperform dedicated translation systems and provide a stronger initialization for reinforcement learning. Our results suggest that fine-tuning is not just about how much a model changes, but how that change is spent.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.