Lost in the Long Tail: Improving Creativity and Coherence in Language Generation
Abstract
Decoding strategies strongly shape the behavior of language models in open-ended generation. Prior work has focused on truncating the distribution over next tokens in order to remove noise. However, do improbable tokens always constitute noise? We examine a phenomenon we term token myopia, in which truncation decoding methods systematically under-generate valid multi-token lexical units because their initial token receives low probability despite being followed by a highly probable continuation. Using idioms and collocations as a measurable testbed, we show that language models often encode idiomatic structure internally yet fail to realize these expressions in free generation due to this first-token bottleneck. To address this limitation, we propose Lookahead-based Rescoring (LoRe), a short-horizon sequence-scoring decoder that rescales candidate next tokens using d-step greedy lookahead continuation scores. To balance quality and diversity, we combine this approach with a stochastic activation mechanism that applies lookahead at a tunable rate. LoRe is complementary to existing truncation-based methods such as min-p sampling, allowing the model to generate more long-tail expressions without altering the underlying model. Across creative generation tasks, LoRe increases multi-word expression usage by 50% on average relative to min- sampling. Furthermore, it outperforms the baseline in 86% of pairwise evaluations on average in creativity, while preserving fluency and coherence. These results suggest that many decoding failures do not stem from deficiencies in model knowledge, but from myopic token-level decision making, and that applying a short-horizon lookahead can better align generation with the latent linguistic structures already represented within language models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.