Guarding a Needle in the Haystack: A Real-Time Policy-Following Streaming Video Guardrail
Abstract
With the rapid growth of video generative models, robust guardrails are more critical than ever to ensure both video content safety, which prevents the proliferation of harmful material (e.g., sexual or self-harm), and video generation security, which defends against adversarial attacks on video generation models (e.g., jailbreak prompts or unsafe video injection). While recent multimodal large language model (MLLM) based guardrails have advanced through reasoning and video understanding, they still face significant limitations. In particular, they rely on frame subsampling, which is unreliable for precise long-term monitoring. In addition, they lack support for real-time streaming and incur high overhead due to inefficient token usage. To address these challenges, we propose StreamGuard, the first real-time, policy-following streaming video guardrail. To precisely identify the few unsafe frames hidden among many benign ones, StreamGuard efficiently inspects the input video in streaming form to localize unsafe content with high precision. To enable real-time streaming, StreamGuard employs an efficient asynchronous inference stack that parallelizes safety analysis across ingested events while simultaneously encoding and detecting incoming frames, achieving fine-grained, frame-level monitoring with low latency. In addition, considering the lack of benchmarks that reflect real-world video safety risks, we introduce two benchmark datasets that include: (1) Safe2Shot, with 5K short videos annotated with frame-level unsafe spans, capturing needle-in-the-haystack cases where harmful content appears in only a few frames; and (2) AdvVideo-Bench, which includes both TV2V and TV2T components targeting the video and text modalities respectively, designed to evaluate guardrail resilience against video-centric multimodal jailbreaks. Extensive experiments show that StreamGuard outperforms the strongest baselines by 7.0 and 6.0 accuracy points on TV2T and TV2V, improves video-level F1 by 1.7 points on Safe2Shot, and achieves the best F1 on four of six existing public benchmarks, while cutting end-to-end latency by 46.1% via parallel event-level inference.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.