Overtopping in LLMs: Spiking-Like Causal Control and the Limits of Static Intervention
Abstract
Unwanted behaviours can emerge without large changes in ordinary model performance, making them difficult to detect through behavioral evaluation alone. We therefore study overtopping, where intervening on a single activation channel affects behaviour across many examples, as a way to examine concentrated causal control inside the model. We ask what kind of causal control these high-leverage channels provide, how they behave under intervention, and whether they remain useful as the model continues to learn. Across five tasks and several model families, we find a recurring spiking-like pattern along two intervention dimensions: how much a channel is perturbed and when during generation the perturbation is applied. As intervention magnitude increases, behaviour can remain stable until a threshold is crossed, then switch and stay changed; across generation steps, the effect tends to be often concentrated at a particular step. We also find strong interactions among high-leverage channels and that the channels carrying high causal leverage can reorganize during learning even when ordinary task performance changes little. As a proof of concept, we introduce a rare trigger-dependent behaviour through controlled poisoning of a grammatical-acceptability task. We then ask whether overtopping measurements can identify useful intervention targets as this behaviour is learned. Channels that suppress the poisoned behaviour at one checkpoint often lose that effect after further training. However, measurements of the current model can still identify single-channel interventions that reduce the poisoned behaviour while limiting disruption to ordinary behaviour. Overtopping therefore reveals localized causal control that is threshold-like, interacting, and able to reorganize during learning. These properties explain both why such channels can provide useful intervention targets and why defenses based on fixed targets may become unreliable as learning continues.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.