acceptodds
Under review as a conference paper at ICLR 2027

Turning Excluded Data into FUEL: Non-Coreset Unlearning Aids Coreset Fine-Tuning in Preserving Model Capabilities

Abstract

Fine-tuning is the dominant approach for adapting large language models (LLMs) to downstream tasks, yet it can degrade previously acquired capabilities, including general knowledge, reasoning, and safety alignment. We show that machine unlearning, conventionally used to remove unwanted data or capabilities, can instead play a constructive role in preserving model capabilities during fine-tuning. Specifically, we study capability control in a target-data-only setting, where the capabilities to be preserved are unknown a priori. We first show that coreset fine-tuning alone is insufficient: even an oracle-selected coreset cannot prevent non-target capability degradation. However, we find that the otherwise discarded non-coreset data can provide a corrective signal when its fine-tuning objective is reversed. Building on this finding, we introduce FUEL (Fine-tuning through Unlearning-Enhanced Learning), a leader–follower framework that fine-tunes on the coreset while unlearning the non-coreset. Across multiple tasks and model scales, FUEL maintains strong target-task performance while substantially reducing degradation on non-target tasks. Moreover, FUEL helps preserve the robustness of safety-aligned models under subsequent benign fine-tuning. Overall, these results reveal a constructive role for unlearning beyond data deletion.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.