acceptodds
Under review as a conference paper at ICLR 2027

PAT: Bending Parameter and Context Scaling Curves through Efficient Context Rewrites

Abstract

Causal full self-attention (FA) caches token representations without revising them when later evidence arrives. This limits efficient storage and relationship composition. We analyze local context rewriting, which converts completed chunks into memory states trained through future-token prediction. These states compose relationships in parallel (A→B and B→C yield A→C) and remove redundant distinctions in the cache, freeing capacity for prediction. Our analysis shows that, under locality assumptions, rewriting reduces the expected layers needed to combine a chain of relationships from linear to logarithmic in its length. The Predictive Abstraction Transformer (PAT) implements these context rewrites within FA’s compute budget by attending to a smaller cache. Its KV cache grows at about one-third FA’s rate, training is fully parallel, and measured training throughput exceeds NanoGPT’s. Experiments show a widening quality advantage over FA as context length and model size grow, while PAT’s relative compute falls with context length. At matched compute within each context length, PAT’s perplexity reduction grows from 4.2% at 2K to 8.0% at 8K. At approximately 97 EFLOPs, 379M PAT achieves lower perplexity than an FA model with 2.3× as many parameters (861M). PAT also improves recall at lengths beyond training: after training on sequences up to 256 tokens, it achieves 95.3% multi-query associative-recall (MQAR) accuracy at 1,024 tokens, versus FA’s 31.9% and 31.7% for KDA.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.