A Harmonic Basis for Auditing Learned Sequential Biases in LLMs
Abstract
In this paper, we show that the next‑token statistics of any autoregressive language model, when observed at a chosen context order , induce a weighted random walk on the de Bruijn graph . We develop a harmonic analysis of this induced walk using the recently derived closed-form eigenvectors of the de Bruijn graph Laplacian. Each eigenvector captures an independent pattern of token co-occurrence at a definite oscillatory scale, and projecting the model's stationary distribution onto this basis yields temperature-independent coefficients – the model's harmonic spectrum – that depend only on the model's centred logits and account exactly for every deviation from uniform generation. We validate the framework with several experiments. First, the harmonic coordinates reveal distinct sequential biases in the task of binary random sequence generation with LLMs. Second, the harmonic spectrum decomposes the exploitable structure of LLM strategies in Rock-Paper-Scissors game into interpretable components, and correlates strongly with win rate across 19 models. Finally, we apply our framework to open-ended generation, projecting part-of-speech tag sequences onto the harmonic basis.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.