When Does Character Training Survive Post-Training?
Abstract
Character midtraining seeks to shape language-model behavior by teaching the principles behind desirable traits before instruction tuning and reinforcement learning. An improvement at this intermediate checkpoint does not establish that the behavior will remain after further training. We investigate this problem through a controlled study of honesty-related behavior in an open language model. We compare models trained on character text with neutral controls matched on corpus size, topic, and reasoning markers, and follow both through instruction tuning and reinforcement learning. By varying whether the midtraining corpora are rehearsed during later stages, we examine how downstream data composition affects behavioral durability. The initial advantage in correcting false premises and acknowledging uncertainty is no longer evident after ordinary instruction tuning. Rehearsing a small amount of character text alongside instruction data yields an advantage that remains after subsequent reinforcement learning. Renewed character supervision during reinforcement learning also improves the character model's absolute behavior after the initial advantage has faded. Its neutral counterpart declines, showing why a larger difference between models is not itself a measure of recovery. Together, these results identify downstream character exposure as a consequential part of the training recipe: rehearsal changes what survives instruction tuning, and the resulting advantage remains after the tested reinforcement-learning stage. Character training should therefore be designed and evaluated together with the further training that follows it.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.