BridgeComp: Prompt Compression via Reasoning-Conditioned Attention and Anchor-Conditioned Policy Optimization
Abstract
The proliferation of large language models in long-context applications has made efficient inference paramount, and prompt compression has emerged as a prominent paradigm to condense context while preserving task-relevant information. However, contemporary compression frameworks introduce semantic distortion and hallucinations by underestimating reasoning-critical tokens or inflating token importance through superficial word-matching. Moreover, answer-level supervision alone provides only indirect guidance for preserving reasoning-critical evidence. To address these challenges, we introduce BridgeComp, a prompt compression framework that estimates token importance from latent tokens that serve as implicit reasoning states. We append learnable latent tokens to the question and use upper-layer attention to isolate deep relational reasoning signals from shallow lexical noise. To align token selection with both answer prediction and compression quality, we train the compressor with reinforcement learning that jointly rewards answer correctness and supporting-evidence preservation. Experiments on question-answering and long-context benchmarks show that BridgeComp achieves state-of-the-art performance. Under a compression setting retaining only 10% of tokens, BridgeComp improves Exact Match by up to 62.44% over the strongest baseline. For long-context compression, BridgeComp achieves up to an 11.75 compression speedup and reduces peak memory usage by up to 39.14%. Our code is available.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.