acceptodds
Under review as a conference paper at ICLR 2027

Twenty-Five Samples to Blind a Guardrail: Stealthy Benign-Side Poisoning of Self-Adapting Safety Filters

Abstract

Safety filters increasingly learn from experience: self-evolving guard pipelines periodically retrain on incoming traffic, and commercial moderation services learn custom categories from the operator's own stream. Because human annotation cannot keep up with volume, these loops label traffic with cheap, weak models. We show that this practice creates an overlooked attack surface: an adversary whose only capability is submitting prompts can poison the loop. The attacker wraps harmful queries in innocuous-looking framings and submits them at volume; the defender's own labelers naturally mislabel a few as safe, and the next fine-tuning step teaches the guard to tolerate that framing style, opening a hole for genuinely harmful requests. Every poisoning label is produced by the defender's pipeline, so the attack is indistinguishable from ordinary traffic. Across five guards from four families and two harmful-query pools, one batch of twenty-five naturally mislabeled submissions collapses or severely erodes framed-request detection on every guard, benign false positives stay at zero on every set we test, and end-to-end attack success against served targets rises from negligible to material. The realistic full-stream window sharpens the operating principle: the mislabel rate, not the dose alone, is the security parameter. Professional-task framings hold the mislabel rate near half, and a few dozen of them blind the guard surgically while every regression metric stays almost unchanged. We contribute methods for both sides: professional-task farming (PTF) and geography-guided farming (GGF) for the attacker, and anchor-calibrated loss filtering (ACLF) for the defender, which removes all of the poison in our runs at a bounded and front-loaded cost in benign data.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.