ForeSift: Speculative Reasoning for Small-to-Large Prompt Compression
Abstract
Long prompts raise target prefill and KV-cache costs, motivating evidence-preserving compression without modifying the target. Small–large compressors let a lightweight draft inspect the source before one target call, but their static context-, question-, or lookahead-conditioned views can miss evidence needs that become clearer as reasoning unfolds—a selection-before-reasoning mismatch. ForeSift addresses this mismatch in one causal draft prefill: it keeps the request-conditioned profile as a Task Anchor, scales a latent reasoning view by its window-level agreement with the anchor, and selects complete sentences under the exact target-tokenizer budget. The method is training-free, requires no draft decoding, and invokes the target once. Across LongBench and RULER- 32K, two model families, and three budgets, ForeSift achieves the best compressed-method macro in all 12 settings, with especially large gains at the 512-token budget and on variable tracing and aggregation. At 32K, a pure-Transformers deployment provides a 1.71× end-to-end speedup and reduces additional allocation beyond the shared target residency by 73% relative to Dense inference.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.