Compression Is Not Evaluation-Neutral: Fixed RAG Compression Can Distort Reader Comparisons
Abstract
Retrieval-augmented generation (RAG) systems often compare readers after a compressor has changed their evidence. This mixes two questions: which complete pipeline works best, and how much of a reader upgrade survives compression. We show that fixed compression can raise average pipeline accuracy while hiding most of a reader upgrade. We hold questions, retrieved candidates, prompts, scoring, and compressed text fixed while comparing 8–20 readers across five question-answering benchmarks and five compression families. HotpotQA and MuSiQue are the main benchmarks, analyzed under a plan fixed in advance. In the 20-reader HotpotQA panel, the lowest- and highest-scoring readers under raw evidence are 31.8 percentage points (pp) apart before compression but only 7.8pp apart on the same stored output from a HotpotQA-trained RECOMP compressor. On both main benchmarks, lower raw-scoring readers gain more under exact match and token-overlap F1. Reader pairs also change order more often under compressed evidence than in raw replays across question halves. Row-level accounting explains how higher average accuracy and smaller upgrades coexist: compression rescues some wrong answers and damages some correct answers, with a different balance for each reader. We release ragscale to audit declared reader upgrades under raw and fixed compressed evidence. Reader evaluations should run this audit before attributing a compressed-only result to the reader.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.