acceptodds
Under review as a conference paper at ICLR 2027

Ordered Top-K UAPs: Benchmarking, Diagnosing, and a Set-then-Order Attack

Abstract

Ordered top- universal adversarial perturbations (UAPs) are a demanding and under-explored form of attack: *can one perturbation impose a prescribed ordered top- list on held-out images, for an ensemble of models?* We benchmark four ordered top- objectives, developed for per-image single-model attacks, in this setting over 3,180 UAPs spanning single models to an 18-model ensemble. The capability is real, but its reach shrinks as the ensemble grows: one perturbation controls ordered lists of up to eight classes on a single model, but only short lists across an ensemble. Fitted to many models, it also transfers: on seven unseen models, foundation-scale ones included, it imposes the attacker's class on up to two-thirds of images on average, and exact ordered pairs on fewer. No objective wins everywhere, and a single attack success rate (ASR) cannot say why. We therefore decompose ordered success into *set membership, ordering given membership, and cross-model agreement*. Ordering is set by how explicitly the loss encodes order (a control that encodes none sits on the chance line), while membership and agreement track attack strength; and each failure, whether under transfer, at large , under adversarial training or against input defences, is one factor collapsing. Acting on the diagnosis, we propose *Set-then-Order*, which establishes membership with a distribution-matching loss and then orders it with hinges, at matched compute. With one schedule frozen across settings, it is never meaningfully worse than the hinge loss alone, and gains substantially where the membership bought in the first phase survives the second.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.