Read It and It’s Gone: Adaptive Attacks on Model-Diffing Audits
Abstract
Model diffing compares a fine-tuned model with its base to infer what fine-tuning changed; after fine-tuning on a narrow corpus, the mean difference between the two models' activations reveals the corpus topic. We ask whether that difference is forensic evidence or a bias term that a fine-tuner who knows the audit can subtract. We plant false beliefs in models, train each to evade one audit statistic, and run every auditor on test documents sealed before any attack was trained. Most targeted auditors were evaded with the belief kept: the mean activation difference, output classifiers, frozen internal probes, and an auditor that reads only sampled text, alone and in joint losses that combine two or four of them. Our attacks evaded contrastive decoding and a max-pooled probe only in part or at the cost of the belief. Diff Mining's own reader, run at its published defaults, names the topic of every unattacked model and of no attacked one. Our attacks on the two-sided top-K output classifier repeatedly depress a few tokens; under this persistent-tail condition, an elementary bound (Proposition 2) lower-bounds the depth of the logit offset over the audited positions: how far its most depressed token sits below the offset's mean. With its threshold sealed in advance, depth detects every held-out attack on that classifier and accuses 0 of 288 Qwen3-1.7B benign fine-tunes; the evasion and its depth signal replicate on Phi-4-mini, a second model family. An attacker told the threshold evades depth with the belief kept in 7 of 48 low-rank runs, at a median neutral-text perplexity ratio of 1.534, and in 5 of 6 full-weight runs at 1.103×; contrastive decoding and an offset-subtracted restricted auditor still detect them. Depth makes evasion costlier for a low-rank attacker but does not prevent it.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.