acceptodds
Under review as a conference paper at ICLR 2027

Information Abundance Paradox: Long Context Training Undermines Parametric Knowledge

Abstract

Large language models are increasingly trained and deployed with long contexts that span documents, code repositories, and interaction histories. This scaling reflects the implicit assumption that longer contexts improve learning by exposing models to richer evidence. We challenge this view by studying how the context window shapes a model's mode of learning, shifting it between parametric internalization and contextualization. We propose the Information Abundance Paradox, which hypothesizes that abundant task-relevant information in the training context can reduce the incentive to encode that information parametrically, thereby increasing reliance on context. In pretraining with long documents, increasing the context window improves language modeling, natural language understanding, and closed-book QA only up to an intermediate optimum, after which performance consistently declines. In supervised fine-tuning more task-relevant train-time context improves performance with supporting context, but reduces robustness when context is absent or misleading. Our analysis suggests that this behavior arises when longer context provides a lower-complexity path to reducing loss. Consistently, context-rich training shifts gradient pressure from feed-forward modules towards attention modules, and causal interventions show that this shift increases reliance on context during inference. Taken together, these findings support the Information Abundance Paradox and challenge the view that near-infinite-context models are the inevitable endpoint of language modeling, even when high-quality long-context data is abundant.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.