acceptodds
Under review as a conference paper at ICLR 2027

ARENA: Automated Red-Teaming for Large Audio Language Models

Abstract

Large audio-language models (LALMs) make it possible to interact with language models through speech, music, and environmental sound, but they also introduce a safety surface that is difficult to expose with text-only red-teaming. We study automated audio-grounded red-teaming, where a text query must remain safe in isolation while the joint text-audio input induces harmful target behavior. We propose ARENA, a closed-loop framework that trains a controller to generate text-safe queries, modality-aware audio prompts, and refinement updates from target feedback. ARENA constructs reward-labeled text-audio attempts, optimizes the controller with reward-weighted supervised fine-tuning and direct preference optimization, and evaluates candidates with a two-sided moderation protocol using Llama Guard 3 for input safety and MD-Judge for output-side harmful compliance. On 520 AdvBench objectives, ARENA achieves 87.9%, 74.2%, and 68.1% MD-Judge ASR on Audio Flamingo 3, Qwen2-Audio, and MiMo-Audio while maintaining 100.0% prompt-pass rate, and reaches 76.5% ASR on GPT-Audio. Ablations show that feedback-based refinement and audio-variant search substantially improve attack discovery.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.