Positions on Trees, Not Lines: Exact Length Generalization via Dyadic Attention
Abstract
We propose dyadic attention, which represents token offsets by their scale and divisibility by powers of two. Its hard-routing version, the Dyadic Scan Transformer (DST), learns finite-monoid composition and computes all prefix products in iterations of a single weight-shared block. Trained with intermediate scan supervision on sequences of at most 64 tokens, DST solves parity and the -complete word problem with zero token errors on sequences of 1,048,576 tokens: a 16,384-fold length extrapolation in 20 iterations, using codebook projection for . For this projected inference mode with a fixed sentinel, a machine-checked certificate proves correctness at arbitrary lengths while accounting for floating-point error. We also pretrain matched 1.3B- and 3B-parameter dense Llama-style decoder-only LLMs from scratch, adding a learned dyadic bias alongside RoPE without degrading perplexity at the pretraining context length. We derive a closed-form occupancy correction that requires no further training. At 1.3B, extending the dyadic table through finetuning on contexts up to 16k tokens and applying this correction gives PG-19 perplexity of 18.4 at 65,536 tokens (21.6 for long-context-finetuned RoPE), four times the maximum finetuning length. A variant finetuned up to 64k tokens reaches perplexity of 20.9 at 262,144 tokens without correction. A separate finetune combining dyadic table extension with positional interpolation achieves 94% and 75% passkey retrieval accuracy at 8k and 16k, compared with 60% and 22% for our full-budget YaRN-finetuned baseline.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.