acceptodds
Under review as a conference paper at ICLR 2027

Mira: Mitigating Pole Geometry Bias in Recurrent Language Models

Abstract

We show that the Transformer's recurrent successors — such as selective state-space models (SSMs), Receptance Weighted Key Value (RWKV), and Retentive Network (RetNet) — can be characterized as infinite-impulse-response (IIR) systems. The zero–pole geometries of these IIR systems characterize the temporal modes they can represent. However, to trade parallelism, the algebraic structures of the parameterizations of these recurrent systems often give rise to poles constrained to the real axis, thereby inducing an architectural bias in their recurrent dynamics. To unify these recurrent architectures and eliminate their pole geometry bias, inspired by the lattice–ladder realization of general IIR systems, we present Mira (Modal Adaptive Recurrent Lattice), a unified family of linear-time recurrent architectures for building large language models (LLMs) with a highly parallel algebraic structure. Each Mira block realizes a higher-order IIR system that represents temporal modes from token streams through a lattice–ladder structure, followed by a feed-forward network (FFN). This realization enables the model to represent the temporal modes associated with both real and complex-conjugate poles. We pretrain Mira-S, a 12-layer small-scale LLM with 0.14B parameters composed entirely of Mira blocks. Throughout the pretraining, Mira-S achieves higher next-token prediction accuracy and lower perplexity than strong recurrent baselines, exhibits stronger context-window generalization beyond the training length, and attains higher throughput with a dedicated CUDA kernel.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.