acceptodds
Under review as a conference paper at ICLR 2027

Trait Hitchhiking in LLM Post-Training

Abstract

Reinforcement learning (RL) post-training is now a standard stage of language-model post-training, and the compute devoted to it is growing rapidly. Whilst it has been shown to improve capabilities by updating the model to maximise a reward function, it can also change behaviours and internal properties that the reward does not explicitly target. We build a mathematical formalism to describe the evolution of traits during RL post-training, and, in doing so, we recover the Price equation from evolutionary biology. From this, we develop a predictive mathematical model for trait hitchhiking: the amplification or suppression of non-targeted traits as training redistributes probability across completions. We evaluate this Price update expression on language models trained with GRPO and show that the Price estimates closely track the sampled trait dynamics. These results identify trait hitchhiking as a mechanism by which RL post-training can silently amplify undesirable traits.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.