Metadata-Aligned Decoding for Accuracy Gains
Abstract
While metadata is commonly used to curate datasets, its potential as a direct decoding signal remains under-explored. We propose LIME (**Li**nguistic **M**etadata **E**mbeddings), a method that enriches token embeddings with linguistic metadata capturing syntax, semantics, and context. LIME substantially improves decoding accuracy across model sizes from 500M to 2B, while introducing only 0.01% additional parameters at negligible compute overhead. Beyond accuracy, LIME improves tokenization and generative task performance. These benefits persist across model scales, and the decoding accuracy gains extend to five typologically diverse languages, where language modeling improvements scale predictably with tokenizer fertility. The gain stems from metadata that anticipates the token to be generated. Supplied directly, as in our guided variant LIME-G, it raises decoding accuracy by over 40% and outperforms constrained decoding in reasoning and arithmetic. Overall, we establish metadata as a powerful interface for controllable generation, demonstrating that inference-time steering can unlock dormant capabilities in reasoning and arithmetic.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.