Lineage Ledgers: Crediting Sources Once, and Only Where They Looked
Abstract
Counting passages is a cheap way to aggregate evidence, and it hands the outcome to whoever copies most, since copying a passage costs nothing while observing the world does not. Confidence here is credited to lineage-distinct sources rather than to deliveries, and each source is credited only along the directions its observation constrains. Truth discovery and pedigree methods instead treat a source as a single vote. A ledger realizes both parts, through an additive union across sources, a max-register join within a source, and a per-source precision matrix. Its credit comes out invariant to redelivery and to arrival order. A coordinated bottom-k sketch bounds the storage, with a finite-sample error bound that holds for scalar and for directional credit alike and does not grow with the number of deliveries. Two pre-registered experiments test the principle. A frozen 7B-parameter reader loses 0.1167 gold-answer recall when one false paragraph is copied sixteen times. Reading from the ledger recovers the whole loss on one fifth of the context, whereas MinHash, the deployed defence, deletes the true paragraph in half of all records when the falsehood arrives first. On real MuSiQue retrieval, support kept per requirement predicts whether retrieved evidence contains every needed fact 0.1535 AUROC better than the same support averaged into a scalar, which sits at chance within every hop count. Inside a learned multi-hop memory the ledger holds its clean accuracy under sixteenfold redelivery, exact or re-chunked, while the undefended memory loses more than a quarter of it, as does every near-duplicate detector once the copies are re-chunked. When only fragments of a source arrive, the ledger matches reading every fragment while crediting one source instead of sixteen.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.