acceptodds
Under review as a conference paper at ICLR 2027

DAPR: Single-Pass Regularization for Capability Retention in LLM Post-Training

Abstract

Continued post-training of large language models via instruction fine-tuning or preference optimization often degrades previously acquired general capabilities. We study local loss geometry as a perspective on this degradation and observe different endpoint Hessian spectra across optimizers. To intervene without an additional perturbed-gradient evaluation, we propose Descent-Anchored Projected Regularization (DAPR) for Adam-family optimizers, which decouples parameter updates into a dual-component architecture: an unmodified primary task-optimization channel driven by Adam moment estimates, and a state-dependent geometric correction regulated by a moving loss center. Above the center, DAPR executes preconditioned descent, and at or below the center, it applies a stabilized projected ascent correction that limits direct cancellation of the Adam direction, combined with an initialization-anchoring penalty that restrains parameter drift. Local two-step analysis identifies conditions under which the correction acts along a weighted gradient-norm penalty, linking its mechanism to Sharpness-Aware Minimization (SAM) while operating strictly within a single forward–backward pass ( AdamW runtime). Across 10k step SFT and DPO on four model configurations scaling from 4B to 14B, DAPR improves capability retention in extended SFT, with lower endpoint task losses on the evaluated 4B/8B models and setting-dependent results under DPO. Finally, we introduce LA-DAPR, which accommodates heterogeneous convergence rates in multilingual tuning via language-specific dynamic centers, consistently improving cross-lingual capability retention.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.