acceptodds
Under review as a conference paper at ICLR 2027

Asymmetric Looped Transformers: Fixed Global Representations for an Iterative Decoder

Abstract

Looped Transformers increase effective depth by repeatedly applying shared lay- ers, but full-stack recurrence does not explicitly distinguish context construction from iterative state refinement. We introduce the Asymmetric Looped Trans- former (AsymLoop), which separates a single-pass causal encoder from a weight- tied recurrent decoder. The encoder produces layer-specific Global key–value caches that remain fixed across iterations; the decoder jointly attends to these caches and Local keys and values derived from its evolving state under a shared softmax. We further adapt existing heavy-tail-guided layerwise learning-rate allo- cation into a Dynamic Learning Rate Schedule (DLRS) for differentiated matrix- wise updates. Experiments on FineWeb-Edu across four nominal model tiers, from 60M to 1B, show lower validation perplexity and higher average zero-shot accuracy across seven downstream tasks than Vanilla Full Loop at matched train- ing budgets and effective depth. AsymLoop reduces perplexity by 0.98 and 0.30 at the 135M and 350M tiers, respectively, and improves the seven-task average from 49.58% to 50.85% at 1B. Ablations support complementary roles for fixed and evolving KV, while showing that advantages over tuned learning rates and strengthened full-loop baselines are not uniform. These single-seed results moti- vate asymmetric information exchange as a design direction for recurrent language models.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.