Are Datasets Reliable? Pitfalls in Evaluating Decomposition-Attack Detection
Abstract
Decomposition attacks against large language models (LLMs) bypass safety measures by splitting a malicious query into multiple sub-queries that appear individually benign. Evaluating and defending against such attacks requires datasets of realistic attack instances. However, existing decomposition attack datasets suffer from several limitations, one of the most critical being that many of the generated sub-query sequences could plausibly arise from a benign query, making them insufficient as genuine attack instances. To address this limitation, we propose the Intent Ambiguity Score (IAS), a metric that quantifies how plausible it is that a given decomposed sequence could also arise from decomposing some benign query. IAS is computed by predicting a benign query whose decomposition is semantically equivalent to a malicious query's sub-sequence, generating that task's own sub-sequence, and then cross-evaluating whether each sub-sequence is a natural decomposition of the other's original query. We formally establish the validity of IAS and compare it against human and LLM-based annotations. We also develop a new decomposition attack dataset generation pipeline that incorporates an IAS-based filtering mechanism. Our dataset is of high quality, achieving a strong attack success rate against closed- and open-weight models while evading existing defense mechanisms.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.