When Do Safety-Relevant Concepts Become Readable in the Training Process of an LLM?
Abstract
Language models in everyday products are expected to avoid toxic, harmful, or stereotyped text. The concepts behind such text take shape during pretraining, so monitoring training requires knowing when each appears. Existing studies follow a concept across saved checkpoints with the raw score of a readout, a method that separates texts with and without the concept using the model's activations: a linear probe or a few sparse autoencoder (SAE) features. However, a raw score cannot tell when the concept became readable, because the two kinds of text differ in their words and the score can be high before any learning. We show that this point in training can be found, and that raw scores miss it. We measure every readout against its own score at the first checkpoint, when the model is still nearly random, and date a concept when this gain stays above a threshold. On 32 concepts across the checkpoints of our primary model, the raw score of a linear probe would date nine to the start of training. We instead find that toxicity emerges first, early in pretraining (by step 1,000). At our threshold, harm categories, which mark the kind of harm a text describes, emerge later and only under the SAE readout, and stereotype (bias) concepts under neither. The date thus depends on the readout, which a single-readout study cannot see. At the final checkpoint, removing a few SAE feature directions reduces generated toxicity. A set of five directions fixed in advance passes checks on benign prompts and output diversity, replicating in every pre-registered combination of three behaviours (toxicity, insult, threat) and two sizes of the primary model (1.4B and 2.8B).
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.