acceptodds
Under review as a conference paper at ICLR 2027

Signal from Noise: Activation Redistribution Improves Reasoning in Large Language Models

Abstract

Motivated by the puzzling observation that prepending long sequences of meaningless tokens can consistently enhance LLM reasoning performance, we investigate the underlying mechanism and propose a more principled alternative. We find that these improvements arise from a redistribution of activations in the MLP layers, where near-zero activations become less frequent while large-magnitude activations increase, effectively suppressing weak signals and promoting more informative representations. Building on this insight, we propose the Activation Redistribution Module (ARM), a lightweight inference-time technique that directly modifies post-nonlinearity activations without altering the input sequence. ARM adaptively shifts near-zero activations outward, implicitly reproducing the benefits of meaningless tokens in a controlled manner. Extensive experiments across diverse benchmarks and model architectures demonstrate that ARM consistently improves reasoning performance while requiring only minimal implementation overhead.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.