Guided by the HINT: Balancing Capability Acquisition and Retention in Offline Agent Fine-Tuning
Abstract
Supervised fine-tuning (SFT) on offline agent trajectories is the standard approach for training specialized tool-using agents, but forcing models to imitate reasoning and actions token by token may harm other capabilities (e.g., general reasoning, tool calling, code generation) of the base model. In this work, we focus on studying how SFT objectives balance the trade-off between acquiring new capabilities while preserving existing ones better on agent data? By comparing several baselines, standard SFT improves the target benchmark while lowering several non-target benchmark scores; meanwhile, simply constraining distributional drift using KL penalty or limiting the update magnitude did not avoid this regression trend. Motivated by recent token-wise adaptive learning objectives, this work proposed Privilege-Guided SFT (PG-SFT) to leverage turn-level information gain of agent trajectories as indicators to adjust supervision strength. PG-SFT yields a more favorable observed trade-off on the evaluated benchmarks, substantially reducing distributional drift and broad capability degradation at the cost of slight degradation on target-task performance. Our findings suggest that balancing the acquisition–retention trade-off depends not only on whether the model is anchored to its base behavior, but also on where and how strongly supervision should depart from that behavior.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.