CompEvidenceBench: Diagnosing Complementary Evidence in Context Compression
Abstract
Budgeted context compression is usually evaluated only by end-to-end answer accuracy. This aggregate score cannot reveal whether failure reflects unavailable evidence, a compressor that drops a jointly useful context, or a reader that does not use retained evidence. We introduce CompEvidenceBench, an auditable benchmark for reader-conditioned complementary-evidence selection failures. For each question, it fixes a budgeted candidate slate containing closed edit quartets, and uses an offline interaction score to audit whether a joint edit is more useful than its constituent edits under a fixed reader. Reference answers, candidate utilities, oracle choices, and interaction labels are never visible to the compressor. On HotpotQA-derived candidate slates at 512 and 1024 tokens, an answer-dependent candidate oracle is 13.18 and 13.04 F1 points above a frozen additive selector overall; on reader-conditioned complementary cases, the gaps are 41.13 and 33.62 points. Yet a three-seed direct-text selector with interaction supervision does not recover this headroom reliably, and a separate frozen pairwise-decision track also yields no stable gain. In contrast, an external generative compressor improves on the complementary subset under both the construction reader and a Mistral cross-reader, although its unknown training overlap prevents a leakage-free ranking claim. CompEvidenceBench therefore separates available opportunity, learned recovery, and external sensitivity: it is an auditable diagnostic testbed rather than evidence for a solved compression method. Our code and dataset are available at https://anonymous.4open.science/r/CompEvidenceBench-2434/README.md.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.