acceptodds
Under review as a conference paper at ICLR 2027

PHASE: Packet-level Hierarchical Adversarial Semantic Exploitation for Targeted Steering of Retrieval-Augmented Diffusion Models

Abstract

Retrieval-augmented diffusion models (RDMs) enhance generation by conditioning the denoising process on a small set of retrieved examples. However, the dependence on external retrieved context introduces additional security risks at the retrieval–generation interface. Existing studies on RDM security mainly focus on manipulating the knowledge base, retriever, or persistent trigger mechanism, while the security implications of the retrieved conditioning packet itself remain insufficiently studied. In this work, we investigate packet-level targeted steering in frozen RDMs and propose PHASE (Packet-level Hierarchical Adversarial Semantic Exploitation), a training-free and gradient-free adversarial framework that operates solely on the conditioning packet. Our central observation is that retrieval-space relevance is not a reliable indicator of generation-time influence: highly target-aligned retrieved examples do not necessarily form the packet that most effectively steers the downstream generation process. Based on this observation, PHASE combines heterogeneous candidate construction, source-aware semantic scoring, structured packet composition, and generation-aware probe selection to identify effective adversarial conditioning under a frozen generator. Extensive experiments across diverse generation settings, semantic evaluators, packet configurations, and retrieval conditions consistently demonstrate the effectiveness and robustness of the proposed framework. Further analyses reveal non-trivial interactions among packet members and a trade-off between retrieval compatibility and downstream steering effectiveness. These results highlight the need to jointly consider retrieval composition and generation-time behavior when evaluating the security of retrieval-augmented generative models.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.