acceptodds
Under review as a conference paper at ICLR 2027

From Unlearnable Outcomes to Decodable Shortcut Codes

Abstract

Unlearnable examples (UEs) add small perturbations to training images so that models trained on them perform poorly on clean data. Low clean accuracy, however, does not reveal what these models learn. We find that shallow networks fit protected training sets more easily than clean ones while generalizing worse. Moving perturbations between images then separates two behaviors: predictions either follow the label associated with the perturbation regardless of the image, or depend on the combination of perturbation and image. The first behavior offers a direct route to unlearnability, because a model that decodes the perturbation needs no semantic features to fit the protected data. We formalize this route as a decodable perturbation shortcut (DPS) and propose a low-complexity shortcut codebook (LCSC), which learns one code per class that randomly initialized networks can identify on changing images. At with crop-and-flip augmentation, LCSC reduces clean accuracy to near chance on four datasets and five architectures; at , its epoch-60 accuracy on CIFAR-10 is . Its factorized variant gives the strongest protection in eight of nine JPEG settings and at adversarial-training radii of and . Both variants are constructed tens to hundreds of times faster than TAP, SEP, and REM.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.