Do Not Regenerate What is in Your Context: Instant Context Reuse via a Pointer Head
Abstract
Large language models increasingly process long-context inputs containing content that may reappear in their outputs, making redundant token-by-token generation a source of decoding overhead. Existing efficient decoding methods, including multi-token prediction and speculative decoding, reduce sequential computation through parallel prediction or draft verification but do not explicitly learn to locate and reuse such content. We propose the Pointer-Generator Large Language Model (PG-LLM), which turns context reuse into a model-native generation action. PG-LLM emits a learned control token when reuse is appropriate and uses an attention-based pointer mechanism to predict the start and end positions of a relevant source span. The selected span is inserted into the output before autoregressive generation resumes. We train PG-LLM at 4B and 30B scales. Across three code editing benchmarks, PG-LLM 30B uses as many autoregressively decoded tokens as its base model on average and improves editing accuracy by 1.64 percentage points, while preserving comparable overall performance across three code generation benchmarks. With a vLLM serving implementation, PG-LLM improves request throughput over the corresponding base model by – at 4B and – at 30B. These results demonstrate that model-native context localization and reuse can reduce autoregressive decoding and deliver practical serving speedups while maintaining comparable overall code generation performance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.