acceptodds
Under review as a conference paper at ICLR 2027

Doing More with Less: Monitoring Generalization During Pretraining via Virtual MoE

Abstract

How to monitor the generalization of dense LLMs as "generalist solvers" during their pretraining? A typical strategy is to evaluate each intermediate checkpoint on benchmarks of downstream tasks, which is costly and offers limited insights into the memorization-to-generalization transition. We propose to track an LLM's generalization by its capability of "doing more with less", i.e., generating diverse, rich outputs by fewer, highly reusable programs. We further discover that such a property can be captured by the pretraining dynamics of model internal states without additional finetuning or benchmark evaluation. Recent studies reveal a strong correlation between the generalization of Mixture-of-Experts (MoE) LLMs and its pathway (expert choices across layers) reusability across samples. Motivated by this, we develop a "virtual MoE" view of layers in dense LLMs by training a CrossCoder as a routing proxy, which enables us to compute and compare pathways between samples. Across the pretraining checkpoints of Pythia (160M-1B), OLMo-1B, and OLMo2-1B, we observe that (1) the edit distance between "virtual" pathways of relevant samples consistently decreases, which indicates their improved reusability; (2) within shared "virtual" pathways, latent representations become more discriminatively structured, suggesting that reusable pathways preserve input-specific distinctions needed for downstream generalization. Moreover, these two trends strongly correlate with generalization performance observed on downstream tasks. Hence, the two metrics form an efficient monitor, which tracks generalization capability and provides a mechanistic interpretation of pretraining.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.