TransDiff: Plug-and-Play Adversarial Defense for No-Reference Image Quality Assessment via Calibrated Score Consensus
Abstract
Adversarial perturbations can substantially alter the predictions of no-reference image quality assessment (NR-IQA) models while remaining nearly imperceptible. Because these models guide image processing and quality-based selection, manipulated scores can compromise downstream decisions without corresponding changes in perceived quality. We introduce TransDiff, a method for estimating a model’s pre-attack score directly from an attacked image. The method evaluates multiple transformed versions of the image to obtain complementary observations related to its original score. Crucially, transformations themselves affect quality predictions. We account for these effects using score-shift distributions calibrated on images without adversarial perturbations, then combine the calibrated observations through distributional agreement. Our theoretical analysis establishes conditions under which this procedure accurately recovers the pre-attack score despite image mismatch, finite calibration and transformation sampling, residual attack effects, and optimization error. TransDiff retains the shape of the calibrated score-shift distributions rather than reducing them to low-order statistics. Experiments on KADID-10k and KonIQ-10k across nine NR-IQA models and white-box and black-box attacks show that TransDiff improves score recovery over test-time image defenses and matched score-recovery baselines. Relative to a matched Gaussian-shift baseline, it reduces aggregate recovery error by 11.5% under white-box attacks and 7.5% under black-box attacks on the nested stress subset. These results demonstrate the value of retaining calibrated score-shift distribution shape for recovering a model’s original prediction.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.