acceptodds
Under review as a conference paper at ICLR 2027

RAQUEL: Robust Evaluation of Machine Unlearning Through Queries over Aligned Databases

Abstract

To address privacy, ethical, and reliability concerns, machine unlearning has emerged as an essential technique for selectively removing information from Large Language Models. However, existing evaluation benchmarks (e.g., TOFU, WMDP, and MUSE) rely on a narrow range of query forms and often fail to test unlearning robustness to semantic paraphrases, complex compositional, and aggregation questions. To tackle this challenge, we introduce RAQUEL, a framework that synthesizes diverse paraphrased, compositional, and aggregation queries by projecting forget/retain knowledge into aligned relational databases. RAQUEL systematically generates diverse database queries and executes them on the aligned databases to determine whether the query should be answered with the forget or retain knowledge. We construct 9,450 QA pairs from TOFU, WMDP, and MUSE to evaluate five unlearning methods on Llama-3.1-8B and Qwen3-8B-Base. Our analysis shows that existing methods often struggle to answer complex questions involving retained knowledge, a degradation that is largely obscured by the original retain questions in existing benchmarks. We further find that the relative ranking of unlearning methods changes when evaluated on RAQUEL-generated datasets, suggesting that evaluation results are sensitive to query diversity and complexity. In addition, training with RAQUEL-generated data improves the robustness of learned knowledge to novel query forms, while subsequent unlearning again changes the relative ordering of methods. We release our code, checkpoints, and datasets.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.