Frequency-Aware Attention: Tokenizer-Induced Self-Information as an Architectural Inductive Bias
Abstract
The performance of Transformer-based large language models (LLMs) is heavily influenced by the power-law token-frequency distribution induced by the tokenizer and corpus, leading to disparate learning dynamics in which LLMs struggle with long-tail tokens. Prior approaches inject token-frequency statistics through data, loss, output-logit, tokenizer, or embedding interventions; we instead make these statistics explicit as an architectural prior within Transformer attention. We introduce Frequency-Aware Attention (FAA), which injects corpus-derived self-information—defined as the negative log-probability of tokens—as an explicit prior into the attention mechanism in the form of attention-logit bias, value scaling, and gating. At the 150M scale, gate-free FAA—using only attention-logit bias and value scaling—reduces perplexity (PPL) by 3.1% and rarest-decile perplexity (D0 PPL) by 7.2% in a representative run, with only 0.043% additional parameters. Full FAA reduces PPL by 4.9% and D0 PPL by 8.9% on average across three seeds, with approximately 4.9% additional parameters. It also improves ARC-E and ARC-C accuracy by 1.89 and 2.05 percentage points, respectively. Our results suggest that making token-frequency statistics architecturally explicit yields a useful inductive bias for language modeling; PPL reductions persist at 450M and 1.2B.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.