acceptodds
Under review as a conference paper at ICLR 2027

Maps of Failures: Observed LLM Vulnerabilities in Behavioral Descriptor Spaces

Abstract

Safety evaluation often reduces a red-teaming run to attack success rate or a set of successful prompts. We instead study the distribution of discovered failures over an explicitly chosen behavioral descriptor space. Our framework uses MAP-Elites to construct a finite-budget archive in which each occupied cell records the highest Alignment Deviation observed for prompts assigned to that cell. The resulting object is a protocol-dependent conditional vulnerability map. Across three full-budget and five reduced-budget model evaluations, we observe up to 72.32% behavioral coverage and up to 370 vulnerabilities. The primary models exhibit different archive-level signatures: broadly distributed observed failures for Llama-3-8B, fragmented high-deviation regions for GPT-OSS-20B, and no observed cell above the selected threshold for GPT-5-Mini. We also measure how the discovered high-deviation cells change under three input-side defenses on Llama-3-8B. The maps, framework, metrics, and anonymized supplementary artifacts are released for evaluation use.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.