acceptodds
Under review as a conference paper at ICLR 2027

Benign Tasks, Unsafe Orders: The Local-to-Global Safety Gap in Continual LLM Fine-Tuning

Abstract

Continual fine-tuning can alter the safety of large language models even when all downstream tasks are benign. We show that this risk is inherently path dependent: applying the same tasks in different orders can lead to substantially different terminal safety outcomes. More importantly, we identify a local-to-global safety gap: state-conditioned measurements can accurately characterize the immediate safety effect of a candidate task, yet even the realized one-step effect can be a poor indicator of safety after the remaining tasks are completed. Thus, reliable local monitoring does not necessarily support reliable long-horizon task-order decisions. Motivated by this observation, we formulate continual task ordering as a terminal path-value estimation problem and introduce P-DC-SCTR, which jointly models the current state, local transition effects, and remaining adaptation path. Experiments across multiple continual fine-tuning settings consistently reveal strong order-dependent safety variation and confirm the local-to-global gap. Path-aware modeling further provides more informative estimates of terminal safety consequences than local signals, while reliable end-to-end ordering remains an open challenge. These findings suggest that safety under continual adaptation should be modeled over training trajectories rather than isolated updates.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.