acceptodds
Under review as a conference paper at ICLR 2027

LooPiT: Input-Space Specialization for Parallel Loop Transformers

Abstract

Weight-shared recurrence improves language modeling without duplicating the backbone, but sequential loops still execute the full recurrent depth. Cross-loop parallelism (CLP) addresses the decode schedule, but leaves the shared-backbone capacity constraint intact. We study loop-specific capacity under CLP, using Parallel Loop Transformers (PLT) as the substrate, without restoring a full set of unique layers. Alongside the established weight-space route of per-loop adapters, we investigate an input-space route: token-indexed tables private to each layer and loop, with shared projections (S-PLE). On a 135M backbone trained for 100B tokens, S-PLE improves geometric-mean perplexity from 27.50 to 22.91, compared with 26.09 for the tested adapters. A 906M study and two-seed checks support the input-over-weight ordering. Rank and width sweeps test how performance changes with added capacity, while channel ablations examine how each loop uses the learned channels. LooPiT uses approximately 45% fewer backbone parameters than depth-doubled dense models across both scales, with approximately 42% fewer total parameters at the larger scale. On a mobile NPU, the input-only configurations retain a graph decode-latency advantage at both scales. The evidence supports input-space specialization as an effective, compatible alternative to weight adaptation within the tested quality–storage–latency trade-offs; prefill efficiency remains an optimization target.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.