From Allow-or-Block to Risk-Adaptive Moderation: A Multi-Expert Safety Guardrail for Large Language Model Services
Abstract
External safety guardrails are increasingly used to protect deployed large language model (LLM) services from harmful and adversarial prompts. Existing guardrails commonly rely on binary filtering, forcing each input into an allow-or-block decision. While effective for risk-unambiguous prompts, this mechanism is less reliable for ambiguous-risk prompts, such as dual-use queries and adversarial jailbreaks, often allowing unsafe content to reach downstream LLMs or incorrectly rejecting benign requests. To address this challenge, we propose RAGuard, a risk-adaptive guardrail that replaces binary filtering with risk-aware routing and specialized mediation modules built on a shared lightweight LLM backbone. Specifically, RAGuard routes each prompt, using intermediate backbone representations, to a safety refusal expert, a risk sanitization expert, or a semantic enhancement expert according to its risk level. The safety refusal expert deterministically rejects clearly harmful prompts. The risk sanitization expert rewrites ambiguous-risk prompts by suppressing unsafe operational details while preserving legitimate intent. The semantic enhancement expert normalizes and enriches clearly benign prompts to improve prompt clarity and downstream response quality. By replacing binary filtering with risk-specific expert adaptation, RAGuard shifts guardrails from simple allow-or-block decisions to risk-aware prompt mediation. Extensive experiments under a guardrail-oriented security–efficiency–utility evaluation framework show that RAGuard achieves strong safety-intervention performance, low hard-rejection rates on benign inputs, and improved downstream response quality, with guardrail-side computational costs evaluated separately.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.