acceptodds
Under review as a conference paper at ICLR 2027

Understanding Negative Learning in Large Language Model Fine-tuning

Abstract

Negative learning has never been more important than it is today, particularly in large language model (LLM) fine-tuning, where it underpins machine unlearning, preference- and reward-based optimization, and exploratory data generation. Yet its functional role remains a subject of ongoing debate, with existing work variously characterizing it as forgetting, regularization, or target shaping. However, analytical support focuses primarily on the single-update scale, while empirical results at the trajectory scale are typically obtained under negative-only training and may not extend to composite objectives in which it interacts with other components. To develop a deeper understanding, we introduce a gradient-based endpoint probing framework that can examine negative updates separately throughout training. Our analysis reveals that negative learning plays different roles across algorithms and training phases. For example, in off-policy DPO, it acts more like a regularizer early in training and shows evidence of target shaping only toward out-of-distribution endpoints and only in later phases. In on-policy PPO and GRPO, by contrast, negative updates offset endpoint-aligned positive learning without stable target shaping, consistent with a role in sustaining exploration. Guided by these findings, we develop a family of stage-aware methods that modulate the strength of negative updates throughout training, improving fine-tuning performance across tasks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.