acceptodds
Under review as a conference paper at ICLR 2027

Homeopathic Learning: Diluting poisoned data can strengthen the behaviour it transmits

Abstract

We uncover a surprising effect in LLM generalization. We find cases where a model fine-tuned only on poisoned examples fails to acquire the target behavior. Yet it does acquire the behavior if we dilute those examples with clean data of the same kind. Paradoxically, adding clean data can strengthen the effect of the poisoned data. We call this homeopathic learning. We demonstrate this in subliminal learning of backdoor behaviors. For example, a student fine-tuned on number sequences from a teacher instructed to respond in German does not acquire this behavior. Yet adding sequences from a normal teacher causes the student to acquire it when triggered by a certain user name. This transfer occurs despite no training example containing German text. Using the same setup, we show that clean data can strengthen the transfer of misalignment. Homeopathic learning extends beyond subliminal transmission. Prior work showed that fine-tuning on archaic bird names can cause a shift to a 19th-century persona. We show that this persona shift is very weak in some models unless the archaic names are diluted with modern ones. We propose a simple regression model as an explanation. Poisoned examples confound the training task with the trigger, while clean examples break this ambiguity, allowing the student to associate the trait with the trigger. Because of homeopathic learning, the same examples can have little effect in one training set and install a misaligned backdoor in another. This exposes a limitation of safety approaches that audit or filter training data in isolation (e.g., one example at a time).

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.