AudioGuard: Toward Comprehensive Audio Safety Protection Across Diverse Threat Models
Abstract
Audio has rapidly become a primary interface for foundation models, powering real-time voice assistants and widely accessible text-to-speech and voice-cloning services. Ensuring safety in these settings is inherently more complex than “unsafe text spoken aloud”: real-world failures can hinge on audio-native harmful sound events, speaker attributes (e.g., child voice), impersonation/voice-cloning misuse, and voice–content compositional harms. These challenges have left the community without standardized definitions, benchmarks, or practical guardrails that capture the full audio risk landscape. To close this gap, we conduct large-scale red teaming of modern audio-capable AI systems and voice generation pipelines, uncovering systematic vulnerabilities unique to the audio modality. Next, we build a comprehensive, policy-grounded audio risk taxonomy and introduce AudioGuardBench, the first benchmark that jointly evaluates audio-input and audio-output safety across diverse threat models, with sliceable coverage over languages, voices (including celebrity/impersonation and child voice), and non-speech sound events. To defend against these threats efficiently, we propose AudioGuard, a unified guardrail that decomposes safety reasoning into complementary components: SoundGuard for waveform-level audio-native cue detection and ContentGuard (ASR + TextGuard) for policy-grounded semantic moderation, composed via an interpretable integration module for scenario-specific decisions. Extensive experiments on AudioGuardBench and five complementary benchmarks show that AudioGuard consistently improves guardrail accuracy over strong audio-LLM-based baselines while substantially reducing latency.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.