A Simpson's Paradox in In-Context Learning Curves
Abstract
A language model's average prediction loss often falls as it reads further into a document. This decrease is commonly used to measure in-context learning, but later positions also contain different tokens. We find a striking reversal on web text: average loss falls while loss rises in every group defined by how often a token has appeared before. Later positions contain more repeats, which have lower loss than first occurrences; their increasing share lowers the average. This is a Simpson's paradox. Across eight text domains and two Pythia sizes, domain accounts for 99% of the variance in a standard position-based score, compared with 0.2% for size. Controlled synthetic experiments also change the score more through the data than through model size. Yet models do benefit from context, and continue improving after this score stabilizes. We therefore recommend reporting token-group losses alongside the average, and measuring context benefit by comparing predictions of the same tokens with different amounts of context. An early loss decrease on first-occurrence tokens provides an additional diagnostic of improvements that the conventional score misses.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.