acceptodds
Under review as a conference paper at ICLR 2027

Are Datasets Reliable? Pitfalls in Evaluating Decomposition-Attack Detection

Abstract

Decomposition attacks against large language models (LLMs) bypass safety measures by splitting a malicious query into multiple sub-queries that appear individually benign. Evaluating and defending against such attacks requires datasets of realistic attack instances. However, existing decomposition attack datasets suffer from several limitations, one of the most critical being that many of the generated sub-query sequences could plausibly arise from a benign query, making them insufficient as genuine attack instances. To address this limitation, we propose the Intent Ambiguity Score (IAS), a metric that quantifies how plausible it is that a given decomposed sequence could also arise from decomposing some benign query. IAS is computed by predicting a benign query whose decomposition is semantically equivalent to a malicious query's sub-sequence, generating that task's own sub-sequence, and then cross-evaluating whether each sub-sequence is a natural decomposition of the other's original query. We formally establish the validity of IAS and compare it against human and LLM-based annotations. We also develop a new decomposition attack dataset generation pipeline that incorporates an IAS-based filtering mechanism. Our dataset is of high quality, achieving a strong attack success rate against closed- and open-weight models while evading existing defense mechanisms.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.