acceptodds
Under review as a conference paper at ICLR 2027

PonderLM-3: Token-Adaptive Pondering with Differentiable Masking

Abstract

Looped language models offer an effective way to improve model performance through iterative computation, motivating a natural follow-up question: when are additional loop iterations actually needed? To address this question, we introduce *PonderLM-3*, a pretraining framework for token-adaptive pondering, where easy tokens stop pondering early while difficult tokens receive additional pondering steps. This allows inference computation to be dynamically allocated at the per-token level. Under purely self-supervised training, PonderLM-3 learns how many pondering steps to allocate to each token through a differentiable attention masking mechanism, which also maintains train–inference consistency. At inference, the learned allocation is realized through hard pruning. The framework builds on the PonderLM-2 backbone. Compared with existing loop-based baselines, PonderLM-3 defines a stronger Pareto frontier, achieving lower pretraining perplexity at equal inference FLOPs. It also achieves the highest average performance on downstream benchmarks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.