Benign Tasks, Unsafe Orders: The Local-to-Global Safety Gap in Continual LLM Fine-Tuning
Abstract
Continual fine-tuning can alter the safety of large language models even when all downstream tasks are benign. We show that this risk is inherently path dependent: applying the same tasks in different orders can lead to substantially different terminal safety outcomes. More importantly, we identify a local-to-global safety gap: state-conditioned measurements can accurately characterize the immediate safety effect of a candidate task, yet even the realized one-step effect can be a poor indicator of safety after the remaining tasks are completed. Thus, reliable local monitoring does not necessarily support reliable long-horizon task-order decisions. Motivated by this observation, we formulate continual task ordering as a terminal path-value estimation problem and introduce P-DC-SCTR, which jointly models the current state, local transition effects, and remaining adaptation path. Experiments across multiple continual fine-tuning settings consistently reveal strong order-dependent safety variation and confirm the local-to-global gap. Path-aware modeling further provides more informative estimates of terminal safety consequences than local signals, while reliable end-to-end ordering remains an open challenge. These findings suggest that safety under continual adaptation should be modeled over training trajectories rather than isolated updates.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.