acceptodds
Under review as a conference paper at ICLR 2027

When Neurons Collide: Why Independently Successful Computations Fail Together

Abstract

Large language models routinely carry out multiple computations within a single forward pass. Because these computations read from and write to the same residual stream, they can interact with one another. We study neural collisions: cases in which computations that succeed independently fail when processed together, even when both share the same answer. Across Llama-3.1-8B-Instruct, Gemma-4-31B-it, and Qwen3-32B, we find that neural collisions are the extreme case of a broader phenomenon: concurrent computations routinely influence one another during inference, sometimes enough to overturn an otherwise correct prediction. In a collision, one computation corrupts another, pushing its internal representation away from the state it reaches in isolation. Using activation patching, we show that this corruption emerges abruptly, at a small number of token positions over just a few layers. To quantify this corruption, we introduce residual drift, the distance between a computation's joint and isolated trajectories. Across all three models, residual drift strongly predicts how much a collision changes the model's answer.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.