acceptodds
Under review as a conference paper at ICLR 2027

OriginCount: Counting Origin Events Instead of Pages in Retrieval Agents

Abstract

A retrieval agent answers a question by fetching pages, counting how many of them support a candidate answer, and stopping once the count is high enough. That count is usually a count of pages. On the web many pages restate one original record, a wire report, a press release, or a dataset card, without adding a new observation; we call such a page an echo of its origin. An echo carries no evidence beyond its origin: if pages are produced from a record and add nothing to it, then , so a page count applies the evidence of once for and once more for every echo. On six web snapshots with page-level provenance labels, of retrieved pages are echoes, and counting by origin instead of by page changes the stopping round or the answer on of the questions. OriginCount groups pages by the origin they can be traced to, gives each group one vote, keeps a page of unknown origin as its own vote, and leaves the content of every page pooled. On a Qwen3-32B host it lowers the rate of confidently wrong answers from to and the calibration error from to , at a cost of one point of accuracy, which the true labels trace to grouping errors, and a rule-based provenance extractor makes it cost ms per page; every comparison is paired by question at one budget of eight prefetched calls. The correction repairs copy inflation and nothing more: it cannot make distinct origins independent, and where copies had propped up a correct answer, grouping them lowers accuracy.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.