Presentation-Order-Gated Sensitivity Implants
Abstract
We construct a small, label-preserving modification of fine-tuning data whose effect depends on the batch order. Under a uniformly random order, the modified data train a model that behaves like one trained on clean data. Under the designated order, which presents the modified batches last, the same data train a model that predicts a prescribed class whenever a trigger pattern appears in its input. This order is not recorded by a fine-tuning run, and a checksum of the data does not cover it. In one pass of mini-batch gradient descent, later batches change the final weights more than earlier ones, and the base model determines how fast this effect decays with position. When the modified batches are presented last, they move the weights toward the implant. When a random order spreads them over the pass, the clean batches that follow them cancel their contribution. The modification of each batch is designed for the position at which that batch appears in training. Experiments on five tasks in images, audio and graphs confirm this behavior: the designated order implants the response, a random order of the same data does not, and the delivered data remain close to clean data at the base weights. The response requires the exact position of each modified batch, and it persists when the designated order is repeated over many epochs.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.