acceptodds
Under review as a conference paper at ICLR 2027

A Dual-Hypothesis Reasoning Framework for LLM Guardrails

Abstract

We propose ARBITER, a novel LLM guardrail framework that introduces two novel ideas: (i) dual-hypothesis reasoning, a reasoning method for LLM guardrails that explicitly reasons over both safe and unsafe interpretations of a prompt before making a safety decision, and (ii) multi-component supervised fine-tuning (MC-SFT), a em structured training loss for reasoning-based guardrails that decomposes LLM outputs into logical components and weighs them by importance. Existing reasoning-based guardrails often rely on expensive procedures, such as generating reasoning traces from larger or closed-source teacher models and applying full-parameter fine-tuning. In contrast, ARBITER uses a cost-effective self-generation strategy for reasoning traces and LoRA-based parameter-efficient fine-tuning, while still achieving better performance than these expensive approaches. Additionally, ARBITER provides faithful evidence-phrase explanations for unsafe decisions, enabling a more transparent and interpretable guardrail method. Experiments on three safety moderation benchmarks show that ARBITER outperforms existing reasoning and non-reasoning guardrail baselines, with clear gains in out-of-domain evaluations.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.