The Cost of a Rule: Why a Backdoor Needs a Count and a Wrong Label Needs a Fraction
Abstract
A network can memorise training examples one by one, or learn a rule: a pattern shared by several examples, mapped to one label and applied to unseen inputs. How many examples a rule needs has two answers that seem to conflict. A backdoor, a trigger pattern added to training images relabelled to a target class, needs a roughly fixed count of these poisoned examples, whereas a class changes its prediction only when about half of it is relabelled, a fraction. We interpret both thresholds as a sum of two costs, counted in examples, and turn this into a cost model. A build cost separates the pattern from competing features, and counter-examples, images with the same pattern under their true label, add a weighted cost that separates it from competing labels. The model gives both regimes as limits: a count when nothing contradicts the pattern, and one half inside a class, whose correctly labelled images are all counter-examples. It also makes four predictions: counter-examples leave the build cost unchanged, the threshold is linear in them, different kinds add, and a backdoor forms at parity, against an equal number of counter-examples, only if each cancels less than one poisoned example. We test these predictions with a small synthetic patch as the pattern. All four hold, and both sides of parity occur, depending on how predictive another feature in the image is, while the pattern moves mainly the build cost. We therefore read the constants as properties of the network's representation, and expect the equation's form to transfer.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.