acceptodds
Under review as a conference paper at ICLR 2027

Silent Loss of Image Conditioning Under Direct Fooling Objectives: A Diagnostic for Amortized Adversarial Attacks

Abstract

Amortized adversarial attacks initially train a model to generate images that can fool a victim classifier, without requiring access to it at attack time. Previous works mention that models trained with an attack objective may collapse and generate similar perturbations across all images, while maintaining high attack success rate and perceptual quality. Our experiments reveal a novel form of this failure where the perturbation remains diverse unlike the earlier failure mode, yet when the generated perturbations are applied to an image other than the one it was generated from, the attack success rate still remains quite high. We introduce a shuffled-image test that compares a perturbation paired with the source image and a non-source image to measure this failure. We also evaluate a partial solution to this problem by guiding the generator with an attacker during an initial warmup phase. This strategy successfully preserves source-image specificity during the warmup phase, keeping perturbations image-specific. However, this strategy cannot consistently keep specificity high after warmup ends and the attack objective starts. Only eight of twelve full-recipe runs kept specificity high to the end. Finally, we test 57 third-party public amortized attack model checkpoints and find high specificity among them, which points to a recipe-specific finding within that one lineage rather than a general defect of amortization.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.