CERTCOMPILE: Preserving Evidence under a Delivery Budget
Abstract
Retrieval agents often read more evidence than they can pass to a downstream reader. Selecting records by relevance alone can omit intermediate facts needed to support an answer. We present CERTCOMPILE, which records source spans and execution dependencies, prioritizes supporting records within a delivery budget, and verifies the evidence before returning an answer. Rejected executions fall back to the original agent. On 120 development questions, a 512-token execution budget and a 256-token delivery budget yield 12 valid accepted answers with dependency-first packing, compared with five for rank-order packing and one for forwarding all admitted evidence. Exact-match accuracy including fallback is 39.17%, 38.33%, and 35.83%, respectively. The policies agree when the budgets are equal. On 240 MuSiQue questions with two models, prioritizing model-generated citation dependencies does not improve answer quality over the same producer's ranking. These results support preserving execution dependencies when execution and delivery budgets differ, while showing that better evidence retention need not translate into comparable gains in answer accuracy or lower inference cost.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.