acceptodds
Under review as a conference paper at ICLR 2027

Context-Demand-Aware Control for N-gram Transformers

Abstract

N-gram embeddings encode local context before attention, but using them only as input features leaves their potential to guide later computation unexplored. We introduce context-demand-aware control for N-gram Transformers, repurposing the N-gram branch as a token-level runtime control signal. A learned certainty score coordinates attention-output gating and embedding recall along depth with attention span and locality bias along the sequence. A shared predictive-reactive controller updates this score using layerwise attention statistics. At 1.3B parameters, the full model reduces held-out perplexity from to across three seeds, a improvement at training-time overhead. A context-removal probe links low certainty to greater dependence on distant context, with the contrast persisting after matching token frequency. Matched-data comparisons and signal-replacement ablations further separate the control signal's contribution from the enriched input representation. Single-seed results at 2.8B and 3.93B support the observed ordering. At 1.3B, optional certainty-driven skipping yields a measured speedup for a PPL increase. These results support using an existing N-gram branch to coordinate computation, within the evaluated N-gram architectures and language-modeling setting.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.