acceptodds
Under review as a conference paper at ICLR 2027

ATMA: Long-Context Language Modeling via Polar Attention and Gated-Delta Compression Memory

Abstract

What determines whether a language model extrapolates beyond its training length? We study attention–memory interactions and checkpoint variability using ATMA, a 378M-parameter hybrid combining Polar Attention with gated-delta recurrent memory. Polar uses a normalized direction, a bounded participation-ratio magnitude, and learned length scaling. In a complete 120-cell, 1B-token factorial, memory improves Polar's teacher-forced target-token retrieval at 64K in all 20 matched cells, by 47.8 percentage points on average. We then train matched NoPE, RoPE, and Polar models for 9.816B tokens at length 2K. The primary Polar checkpoint retains 62.7% exact FinePDFs retrieval at 8K and 34.4% target-token accuracy and 9.0% exact five-token accuracy overall at 256K, with all exact successes on synthetic haystacks. Two fresh paired runs retain a retrieval advantage over NoPE but reach only 0.0% and 3.83% exact accuracy. A post-hoc intervention caps the most retentive head in each checkpoint: primary NoPE and Polar 256K bits per byte fall from 8.097 to 1.595 and from 1.825 to 1.501. Both NoPE replications recover the likelihood improvement; the Polar replications have no comparable gate outlier and remain essentially unchanged by the cap. A separately optimized TDA hybrid reaches 0.2% token retrieval but 40% adapted BABILong accuracy at 256K, exposing distinct retrieval and reasoning trade-offs. Prospective controls resolve attribution: across three matched initialization triplets, Polar strictly outperforms temperature-matched softmax on natural-text exact retrieval by an average of 22.18 percentage points, and retains a 37.18-point advantage across 30 independent documents. The one-head retention cap rescues likelihood without conferring a general retrieval boost, delineating the boundaries of retention stability.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.