acceptodds
Under review as a conference paper at ICLR 2027

Adversarial Examples, Inverted!

Abstract

Almost all adversarial attacks add an imperceptible perturbation to fool a model. We instead study the opposite: a large, clearly visible perturbation that causes the model to keep its original, correct prediction, even though a human would no longer recognize the image. We show such examples can be generated at scale, and then answer three questions this formulation leaves open: whether humans really perform worse than the model, whether standard out-of-distribution (OOD) detection and calibration tools catch it, and whether existing defenses mitigate it. We answer all three on MNIST, CIFAR-10, and ImageNet, and additionally verify that none of these answers depends on the specific attack variant or source architecture used to generate the stimuli.Throughout, we use NI-FGM (momentum-free, -normalized) as the primary attack and VGG-11 as the primary CIFAR-10 architecture, with a dense grid (12 points, up to 22 on CIFAR-10 and 13 on MNIST) and samples per track. Sec. sec:res-robustness reports a robustness check in which the entire pipeline is instead run with the momentum-based NMI-FGM attack and ResNet-18 as the source architecture, confirming every conclusion below is unchanged. (i) An independent recognizer proxy falls from near-ceiling to 20% on CIFAR-10 by the top of the widened range while the model stays at 100% — an 80-point gap, larger than under the original configuration, corroborated directly by a small human pilot () and not explained by signal loss (a matched-magnitude Gaussian control degrades recognizability faster); a CLIP zero-shot proxy confirms the gap at ImageNet scale too. (ii) Confidence- and energy-based OOD detectors and calibration are structurally blind (0% detection, ECE ), while a feature-space Mahalanobis detector flags 100% — but is evaded by an adaptive attacker at no cost to success. (iii) No classical defense, including adversarial training (45% robust accuracy), reduces attack success (correlation with large- resistance ). A mechanistic analysis further shows the attack destroys low-level texture far faster than edge/shape structure. Code, figures, and a ready-to-run human-study harness are included as supplementary material.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.