acceptodds
Under review as a conference paper at ICLR 2027

AAPO: Adaptive Adversarial Prompt Orchestration for Robust Unlearning in Large Language Models

Abstract

Machine unlearning aims to selectively remove specific knowledge from large language models (LLMs) to meet privacy, safety, and compliance requirements. However, existing unlearning methods are primarily evaluated under non-adversarial settings and often fail to maintain the unlearning state when faced with adaptive adversarial prompt attacks at inference time. This fragility of the unlearning state severely undermines the practical value of machine unlearning and may pose latent risks for real-world deployment. Existing studies have attempted to introduce defensive mechanisms to improve model robustness. Nevertheless, the majority of current approaches are dependent on parameter updates, which not only introduce computational overhead but also significantly limit the generalization capability to diverse unknown attacks in the future. To address these challenges, we propose Adaptive Adversarial Prompt Orchestration(AAPO), a model-decoupled, inference-stage defense framework. AAPO formulates robust unlearning as a closed-loop adversarial procedure involving Attacker, Judge, and Defender agents. It enables the dynamic evolution of defense prompts via multi-agent self-play to adaptively address a wide range of continuously evolving potential attacks. Extensive experiments on the RWKU benchmark across multiple mainstream LLMs demonstrate that AAPO significantly improves unlearning robustness under adversarial attacks, achieves up to a 70% relative reduction in forget-set leakage across multiple baseline methods, while preserving overall model utility. Additional experiments on the WMDP benchmark further validate the generalizability of AAPO across different unlearning settings.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.