acceptodds
Under review as a conference paper at ICLR 2027

GuardAlign: Guardrail-Aware Safety Alignment for Modern LLM Pipelines

Abstract

External guardrails increasingly complement aligned LLMs against jailbreaks, yet guardrails and backend models are typically optimized independently. Consequently, an unsafe prompt missed by an input guardrail reaches a backend never explicitly aligned to the guardrail's coverage gaps. We introduce Guardrail-Aware Safety Alignment (GuardAlign), a paradigm that retains the full safety preference dataset and augments each example with offline decisions and normalized unsafe-confidence evidence from multiple input guardrails. Within GuardAlign, Guardrail-Aware Preference Optimization (GuardPO) is a reference-free objective that replaces a static margin with an example-specific target margin. Its Detection Penalty emphasizes guardrail misses and low-confidence detections, while its Stubbornness Penalty targets current-policy misrankings only when safety and helpfulness agree. Across four LLM backbones and six safety test sets, every GuardAlign variant attains a lower overall mean pipeline attack success rate (ASR) than Pre-Guardrail, Post-Guardrail, and Self-Refine; GuardPO performs best at 2.42%, versus 7.23% for the strongest of these baselines. Without an inference-time guardrail, GuardPO also reduces adaptive-attack ASR to 23.75%, compared with 31.00% for G-SimPO and 43.75% for G-DPO. It further incurs substantially less online overhead than output filtering and iterative refinement, while retaining the most BIG-Bench Hard utility among the evaluated guardrail-aware objectives.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.