Detectability Inversely Tracks Disruption: Reference-Free, Static-Weight Detection of Backdoored LoRA Adapters
Abstract
Can a backdoored LoRA adapter be screened without running its base model? We study directional statistics computed from adapter weights alone. Our raw circuit_score is a nonnegative ratio of input-output alignment to cross-layer output coherence; its common normalization cancels without a benign-reference assumption. Across three language-model architectures, two ranks, and 15 seeds, a separate supervised axis scorer achieves AUROC 0.96-1.00 in five of six settings; the raw score is weaker and reverses its ordering in one setting. Across three vision-language backdoor designs, supervised detectability decreases as behavioral disruption increases: the surgical design reaches AUROC 0.96 (0.90 with benign-only calibration), despite 92.6% attack success and near-zero clean-accuracy loss. A constrained attack-defense formulation distinguishes behavioral stealth from weight-space concealment. Adaptive training further shows that suppressing one statistic can preserve attack success while increasing exposure to an untargeted readout. These experiments establish competing objectives in the tested designs, rather than a universal tradeoff or an equilibrium. We distinguish reference-free feature computation, supervised detection, and benign-calibrated screening, and identify the calibration and generalization limits that prevent interpreting ranking performance as deployment assurance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.