acceptodds
Under review as a conference paper at ICLR 2027

CovDPO: Bridging Representation-Space Coverage Gaps in LLM Safety Alignment

Abstract

Despite their effectiveness on in-distribution requests, safety-aligned models often generalize poorly to out-of-distribution (OOD) harmful inputs. Across two model families, we find that harmful inputs associated with safety failures tend to lie in regions weakly covered by the training safety-preference data, indicating that safety generalization failures are systematically associated with gaps in representation-space coverage. We therefore propose CovDPO, a coverage-aware framework that connects coverage diagnosis with preference data augmentation and optimization. CovDPO introduces the Hidden-State Coverage Score (HCS) to measure the local coverage of safety-preference data in representation space. HCS serves two complementary roles: it guides preference-data augmentation toward under-covered regions associated with safety failures, and it provides coverage-dependent loss scaling for the DPO objective. Across four OOD harmful benchmarks and three jailbreak attacks, CovDPO outperforms the strongest baseline on both model families, reducing the average attack success rate from 15.90% to 2.41% on Llama-3-8B and from 11.21% to 4.60% on Qwen3-8B, while maintaining low over-refusal.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.