acceptodds
Under review as a conference paper at ICLR 2027

Long-Tailed Entity Distributions Preserve World Model Plasticity

Abstract

Our world is ever-changing and, while language models can incorporate new information flexibly in-context, they face difficulties doing so *in weights*, in a way that generalizes. We show that the pre-training data distribution is a key lever: long-tailed entity distributions preserve a model's ability to integrate new entities into its world model. We study this in a controllable synthetic world where cities with known geometry underlie multiple relational tasks. We first show that models use the representational geometry of entities when solving tasks, and that how well a new entity is integrated into the world geometry predicts its cross-task generalization (). We then find that under uniform pre-training, the world representation forms early but loses plasticity: its ability to integrate new entities via gradient descent decays with continued training, even as pre-training loss decreases. Pre-training with Zipfian entity frequencies counters this, improving new-entity integration on average and on the hardest task, and outperforming 15 alternatives including meta-learning. We hypothesize that sparse exposure keeps some entities effectively new (i.e. incompletely learned), providing continued practice integrating them throughout pre-training. At high Zipf , however, tail learning and world-model quality deteriorate. In all, we establish that long-tailed entity distributions help maintain the plasticity of entity representations and reveal a trade-off between forming a world model and keeping it open to change.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.