acceptodds
Under review as a conference paper at ICLR 2027

Unanticipated Preference Drift: Training Data Outweighs Tasks and Defies Self-Prediction

Abstract

Post-training can change a model in ways beyond the intended objective. Models also have measurable preferences, and training that deliberately targets these preferences can shift them. How preferences shift when training does not target them, and whether models can anticipate these shifts, remain largely unexamined. We measure how coding-task, value, and training-data preferences shift when training Qwen3.5 models (4B and 9B, with selected 27B runs) with coding SFT, Rust reinforcement learning (RL), and SFT on conversations high or low in a personality trait. Before training, we also ask each model how it expects training to affect its preferences. Rust RL barely moves any preferences. Coding SFT moves training-data preferences at 4B but almost no coding-task preferences, while personality conversations shift preferences across coding tasks, values, and training data, moving 19–22 of 27 coding tasks. The direction of these shifts depends on whether the conversations were high or low in a trait, not on which trait. Models anticipate the general direction in which personality training shifts value preferences, but a rule based only on the training description does as well, and no model identifies which coding tasks would shift. Because preferences partly predict behavior, shifts from training need to be measured directly, not inferred from the training data or the model’s forecast. As models increasingly help train themselves and create their own training data, a model that cannot anticipate its own preference drift is limited in steering its preferences through training. It also cannot flag unwanted shifts and may expect shifts that do not happen.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.