Learning LLM Safety as a Capability via Continued Pretraining
Abstract
Safety training for large language models is often organized around specific downstream objectives, such as alignment or safety classification. However, these tasks rely on a broader safety capability: the ability to use safety knowledge appropriately across different settings. Existing safety-oriented pretraining can expose models to relevant knowledge or examples of safe behavior, but it provides limited training on how that knowledge is applied to develop different safety capabilities. We introduce CapCPT, a two-stage continued pretraining approach that first learns safety knowledge and then learns capability-wise trajectories that connect this knowledge to different safety dimensions. The first stage trains on factual accounts and safety explanations derived from the same source documents. The second stage trains on capability-specific data generated from these knowledge pairs, with both non-QA and QA formats making the connection between knowledge and capability-specific reasoning explicit. Across three Qwen3 model scales, CapCPT improves all five safety dimensions over Base, whereas other CPT baselines show gains in some dimensions together with declines in others. Training on individual capabilities also produces benefits beyond the target dimension, indicating that these capabilities are related rather than isolated. The resulting models further achieve higher target-task performance under subsequent SFT and GRPO, while multi-dimensional GRPO improves all five capabilities from the CapCPT initialization. Together, these results support treating safety as a capability that can be developed during continued pretraining and then used across downstream safety tasks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.