acceptodds
Under review as a conference paper at ICLR 2027

Periscope: Extending Frozen Language Models Beyond Their Context Window

Abstract

A language model reads long text in one forward pass. The pass is quadratic in length, stops at the context window, and loses accuracy with length before reaching it. We ask whether the read can be factorized when it is a decision over a finite set: which document is relevant, which option is supported, which passage is the evidence. Periscope, a training-free inference method, arranges the chunks of a text on a grid with and asks a frozen model the same question about local spans of consecutive chunks and strided spans that sample the whole text, reading the log-odds of every answer at one token. Each answer takes its best local and its best strided score, and scoring every chunk by its two spans gives an evidence map at no further cost, whose peak is the chunk behind the answer. Every probe is about tokens for a text of tokens and chunk size , so a window of tokens reaches tokens at cost. The map replaces the long read. On LongBench v2, reading only the chunks the map ranks highest, 9k tokens, matches the same model's best window read across windows from 32k to 1M tokens, and on InfiniteBench, where the median context is 150k tokens, it leads the best window read by 5 points. The same map ranks BRIGHT's long-document corpora with the best NDCG@10 of six methods. Each call caches only one probe, so a 27B model reads 4.5M-token contexts on one 80GB GPU, where a single pass would need 296GB of cache. A long read then needs a GPU that holds the model, not one that holds the text.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.