SCOR: Sequential Counterfactual Ordering for Budgeted LLM Backdoor Suppression
Abstract
Backdoor suppression in a shipped language model is governed by intervention order: component effects interact, every removal spends clean utility, and partial pruning can reroute attack behavior through surviving units. We introduce Sequential Counterfactual Ordering with Remeasurement (SCOR), which measures attention heads, MLP groups, and LoRA ranks through paired clean/triggered interventions, admits each removal by its observed joint utility cost, and re-estimates the ordering after migration. With candidates, pruning, repair, and the 3 pp calibration criterion fixed, SCOR reaches 11.6/41.2% token/semantic ASR, improving on attack-only ordering by 5.6/8.3 pp while recovering 2.3 pp clean accuracy. Across the reported trigger-family cells on Llama-3.1-70B and Qwen-2.5-72B, SCOR attains the lowest ASR among the compared post-hoc interventions while retaining clean accuracy near the undefended checkpoint. These results identify observed-feasible, dynamically remeasured ordering as a practical design principle for shipped-model backdoor intervention.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.