acceptodds
Under review as a conference paper at ICLR 2027

CoverSynth: Coverage-Guided Generation of Retrieval Evaluation Sets

Abstract

Enterprise agents need retrieval evaluation sets that expose domain-specific failures and guide developers towards efficient deployment tradeoffs. Neither expert-human-authored sets nor current synthetic generation methods provide a methodology for achieving the necessary coverage, reliability, efficiency, and discriminability required for such decision-making. Simultaneously, the field does not have an assessment framework for determining when an evaluation set has satisfied these properties. We introduce an evaluation set assessment framework for these criteria and CoverSynth - a generator specifically designed to satisfy these criteria without incurring the high cost of human expert labeling. Starting from an unannotated corpus, CoverSynth targets coverage across retrieval demands and system capabilities. Deterministic checks locate evidence quotes in their sources, while model-based checks assess semantic validity to guarantee reliable ground truth. Across five public enterprise-style corpora, our method provides informative variation in difficulty, revealing unseen failure modes such as an inability of agents to handle underspecified or ambiguous requests. On the RepLiQA corpus, 91% of human-generated queries can be answered by all members on a panel of 12 retrieval systems, compared to only 12% of CoverSynth-generated queries. Finally, in a repeated TechQA/RepLiQA study, 12 of 12 within-harness model contrasts have positive nominal 95% intervals on generated sets, compared with 4 of 12 on human sets. CoverSynth makes coverage and evidence quality explicit construction objectives, producing corpus-grounded evaluations that reveal retrieval headroom and differences between agent configurations.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.