ConSieve: Constraint-Sensitive Evidence Selection for RAG Context Compression
Abstract
Post-retrieval context compressors for retrieval-augmented generation (RAG) typically select evidence based on its semantic relevance to the query. Relevance alone does not guarantee validity because text can remain topically relevant while violating an answer-critical condition, especially under temporal constraints, and therefore fail to support the answer. To measure this failure directly, we construct fixed-evidence counterfactual diagnostic pairs by changing one answer-critical query condition while holding the document and candidate evidence fixed, so evidence that supports the original query becomes invalidated for the edited query. On this diagnostic, the tested baselines retain 55.3% to 72.1% of the invalidated evidence. We introduce ConSieve, a constraint-sensitive extractive compressor trained with counterfactual supervision. On the predominantly temporal diagnostic, ConSieve retains more supporting evidence under the original query than every tested baseline and prunes 87.2% of the same evidence after the edit makes it invalid. A matched ablation links this behavior to counterfactual supervision, reducing invalid retention from 56.1% to 11.4%. Under a shared 256-token context cap in end-to-end RAG, ConSieve reaches 0.631 downstream QA F1 compared with 0.608 for the strongest tested baseline. A separate multi-budget analysis shows the same overall advantage across the tested context budgets.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.