acceptodds
Under review as a conference paper at ICLR 2027

Loss Smoothing for Stable Adaptation Under Distribution Shift

Abstract

Over their life cycles, neural networks are often adapted under distribution shift. Standard adaptation methods typically optimize the target objective directly, inducing an abrupt change from the source training objective. This abrupt change can hurt optimization or distort learned representations, including features that may still be useful for the new task. We investigate how softening this transition can improve adaptation. Specifically, we formulate *loss smoothing* as a simple approach that linearly interpolates between the source and target objectives. We show that retaining the source supervision can improve target adaptation even when preserving source-task performance is not a goal. A controlled toy study shows how smoothing sustains learning of features shared between source and target tasks. In pretrained vision and language models, smoothing limits the loss of source information during adaptation, with retention advantages that persist after source supervision ends. Across a broad variety of adaptation settings, including pretrained vision model adaptation, continual supervised learning, fine-tuning "overtrained" large language models (LLMs), and offline-to-online reinforcement learning, loss smoothing improves performance, suggesting that smoother objective transitions are a broadly useful tool for model adaptation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.