acceptodds
Under review as a conference paper at ICLR 2027

Learning to Copy, Learning to Depend: Language Models Offload In-Context Copying to a Pointer Head

Abstract

Many next words are copies: a character's name, a number, a term defined a paragraph earlier. We train language models jointly with a pointer-copy head that copies such tokens from context, then remove it at inference: accuracy on LAMBADA, where most targets appear earlier in the passage, collapses below 7%, far below the 18–23% of baselines trained without it, while eight other benchmarks barely move. We find this output-pathway offloading, in which the backbone's vocabulary readout delegates in-context copying to the head rather than duplicating it, in all five architecture families: two Transformers, Mamba-1, a Jamba-style hybrid, and Gated DeltaNet, each trained from scratch as a controlled baseline/+copy pair at 130M–198M, with further pairs from 45M to 350M. In both Transformers, a linear probe decodes copy targets from the +copy backbone's final hidden state as well as from the baseline's, while its native vocabulary readout becomes less accurate on them. With the head in place, WikiText-103 perplexity and LAMBADA accuracy improve in every pair, and held-out perplexity in all five 5B-token pairs. Copyability stratification confines every pair's LAMBADA gain to targets present in the preceding context, and randomizing the retrieved identities drops every model below its baseline. Under one fixed head configuration, Mamba-3 SISO - a backbone redesigned for retrieval - still benefits on every headline metric. A three-seed replication at 60M reproduces every headline gain's sign and the closed-gate collapse. Because the head configuration differs between architecture classes, we do not rank architectures; all 130M–350M results are single-seed.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.