CovDPO: Bridging Representation-Space Coverage Gaps in LLM Safety Alignment
Abstract
Despite their effectiveness on in-distribution requests, safety-aligned models often generalize poorly to out-of-distribution (OOD) harmful inputs. Across two model families, we find that harmful inputs associated with safety failures tend to lie in regions weakly covered by the training safety-preference data, indicating that safety generalization failures are systematically associated with gaps in representation-space coverage. We therefore propose CovDPO, a coverage-aware framework that connects coverage diagnosis with preference data augmentation and optimization. CovDPO introduces the Hidden-State Coverage Score (HCS) to measure the local coverage of safety-preference data in representation space. HCS serves two complementary roles: it guides preference-data augmentation toward under-covered regions associated with safety failures, and it provides coverage-dependent loss scaling for the DPO objective. Across four OOD harmful benchmarks and three jailbreak attacks, CovDPO outperforms the strongest baseline on both model families, reducing the average attack success rate from 15.90% to 2.41% on Llama-3-8B and from 11.21% to 4.60% on Qwen3-8B, while maintaining low over-refusal.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.