acceptodds
Under review as a conference paper at ICLR 2027

Sparks of Generalization: Capacity Windows and Steering During Transformer Pretraining

Abstract

Can a transformer learn rules while retaining facts under the same next-token objective? We study Qwen3 and Gemma3 trained from scratch on controlled corpora with native tokenizers, one head per layer, and one or two layers in both architectures (three in Qwen3). Single-digit addition and subtraction probe in-distribution generalization, requiring all 196 training and four withheld equations correct at the final evaluation. With facts included, success also requires recall of at least 48 of 50 facts at that evaluation (joint success). In one-layer arithmetic-only baselines of both architectures, generalization is most frequent within a limited hidden-width range. At the widest tested width, all 10 seeds per architecture fit the training equations and none generalizes. To steer generalization, we introduce answer consistency, an auxiliary loss on arithmetic representations grouped by operator and training answer. Across thirteen steering blocks (three arms on shared seeds) using arithmetic alone or adding facts and approximately 1,400 prose tokens, it raises success over the no-penalty baseline in nine and over a per-step shuffled-label control in five (nominal, unadjusted paired , all assigned seeds). In one-layer arithmetic-only models at the widest tested width, steering yields no detectable gain in success. With facts and this prose corpus, steered one-layer width-14 models reach joint success in 34/60 Qwen3 and 46/60 Gemma3 seeds, versus 20/60 and 32/60 for baseline and 17/60 and 33/60 for shuffled labels. Steering also tightens same-answer representations at the regularized site across the architectures and depths studied.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.