acceptodds
Under review as a conference paper at ICLR 2027

Dr. Post-Training: A Data Regularization Perspective on LLM Post-Training

Abstract

LLM post-training faces a critical challenge: effectively leveraging scarce, high-fidelity target data alongside abundant but imperfectly aligned general training data, a problem that existing methods frame as data selection. We introduce **Dr. Post-Training** (**D**ata-**R**egularized Post-Training), a framework that instead treats general training data as a data-induced regularizer that stabilizes learning from the scarce target objective. Specifically, our framework proposes that at each training step, construct a feasible set of model update directions that depends on the general training data alone, and project the model update direction specified by the scarce target data onto that feasible set. Standard training and existing data selection methods arise as special cases with different choices of the data-induced regularizer, and these methods correspond to different points on a bias-variance spectrum with different regularization strength. The framework thus provides a principled lens for understanding these tradeoffs and guiding the design of new methods. Building on this view, we propose a family of methods that expose the regularization strength as an explicit design choice, and a theory that ties this choice to how well the general data covers the target signals and how much target noise it lets through. For practical LLM-scale use, we introduce careful system optimizations that realize these methods at a per-step cost on par with standard training. Experiments across SFT, RLHF, and RLVR follow the predicted tradeoff: the best regularization strength tracks the mismatch between the general and target data, with our methods gaining most where the target is diluted in the general data and standard training remaining competitive where the two are aligned.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.