acceptodds
Under review as a conference paper at ICLR 2027

ForeSift: Speculative Reasoning for Small-to-Large Prompt Compression

Abstract

Long prompts raise target prefill and KV-cache costs, motivating evidence-preserving compression without modifying the target. Small–large compressors let a lightweight draft inspect the source before one target call, but their static context-, question-, or lookahead-conditioned views can miss evidence needs that become clearer as reasoning unfolds—a selection-before-reasoning mismatch. ForeSift addresses this mismatch in one causal draft prefill: it keeps the request-conditioned profile as a Task Anchor, scales a latent reasoning view by its window-level agreement with the anchor, and selects complete sentences under the exact target-tokenizer budget. The method is training-free, requires no draft decoding, and invokes the target once. Across LongBench and RULER- 32K, two model families, and three budgets, ForeSift achieves the best compressed-method macro in all 12 settings, with especially large gains at the 512-token budget and on variable tracing and aggregation. At 32K, a pure-Transformers deployment provides a 1.71× end-to-end speedup and reduces additional allocation beyond the shared target residency by 73% relative to Dense inference.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.