acceptodds
Under review as a conference paper at ICLR 2027

Lost in the Gaps: How Non-Lexical Filler Disrupts Belief Tracking in Language Models

Abstract

Unlike natural-language filler, repetitive, meaningless tokens such as a sequence of spaces can significantly degrade Language Models' (LMs) task accuracy when inserted into a prompt. Why does this happen? We investigate this question using a belief tracking task, in which a model must infer an agent's belief about the world, as a case study. Through activation patching and causal interventions on attention across three LMs, we find that filler sequences that are associated with performance degradation receive markedly high attention weights. Activation patching traces these failures to early MLP layers processing the filler: exchanging their outputs between harmful and harmless fillers suffices to remove or induce the errors. Patching individual MLP neurons further localizes these effects to a single neuron in Llama-3.1-8B and four neurons in Qwen2.5-14B, whereas in Qwen2.5-7B a few neurons suffice to remove the errors but inducing them requires a more distributed set. Our findings connect filler-induced failures to early MLP computations and downstream attention interactions, providing a mechanistic account of how semantically uninformative inputs disrupt belief tracking.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.