User Message Fine-tuning: Implanting Beliefs, Preferences, and Alignment in LLMs by Training on User Messages
Abstract
Chat models are typically fine-tuned only on assistant turns, with all other tokens masked from the loss. We study the reverse: User Message Fine-tuning (UMF), a class of fine-tuning techniques in which we train a model on synthetic user messages. Across four experiments on Qwen3.6-35B-A3B and Qwen3-8B, we find that UMF shapes the assistant despite never training on its outputs. (1) We fine-tune a model on a corpus of user messages that treat a false fact as true, and find that this makes the model believe the false fact more deeply than synthetic document fine-tuning (SDF) does, as measured by behavioral evaluations and truth probing. (2) We fine-tune a model on a corpus of user messages that reference a made-up fact about the user, and find that this makes the model believe that the fact is true for new users, even in fresh contexts. (3) We fine-tune a model on user reactions that praise one answer to a preference question and criticize another, and find that this shifts the model's own answers toward the praised option. (4) We fine-tune a model on user messages that react positively to narrowly misaligned assistant responses, and find that this model shows less emergent misalignment after it is later fine-tuned on those narrowly misaligned responses. Together, these results support UMF as a promising new tool for shaping a model's beliefs, preferences, and alignment, with many applications to AI safety such as building model organisms and mitigating unintended generalization.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.