acceptodds
Under review as a conference paper at ICLR 2027

HarmonyGuard: A Multi-Agent Guardrail for Web Agents via Dual-Timescale Policy Grounding

Abstract

Large language models enable web agents to autonomously complete complex tasks in open environments, but untrusted content can redirect benign trajectories toward policy violations. Safety policies state high-level requirements but cannot enumerate how violations emerge in concrete trajectories. This mismatch creates a policy-to-trajectory grounding gap. Runtime violations bridge this gap by showing how abstract policies are breached in context. Existing guardrails use this evidence either for current correction or later judgment, limiting joint gains in task utility and policy compliance. We therefore introduce HarmonyGuard, a multi-agent guardrail based on dual-timescale policy grounding. Under this formulation, the same policy-grounded violation repairs the active trajectory and grounds later judgments. A Utility Agent evaluates adjacent-step transitions and guides policy-compliant revisions, while a Policy Agent structures trusted policies and retains selected violations as bounded, policy-indexed boundary cases. On three agent safety benchmarks, HarmonyGuard attains the highest compliance and completion under policy without suppressing task progress, exceeding the prompt-based baseline by up to 17.8% and 23.8%. It leads the agent guardrail by 8.2% in accuracy and matches trained guardrail models at the lowest false positive rate.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.