When Repetition Becomes Evidence: Semantic Evidence Overcounting in Multi-Agent RAG
Abstract
Multi-agent retrieval-augmented generation (RAG) systems may treat agreement as independent evidence even when reports share a source. We study this behavioral-level mechanism in MADAM-RAG: when content derived from a single wrong semantic unit is copied or paraphrased across multiple surface-level documents, it produces apparently independent agent reports that may be counted repeatedly. We distinguish independent wrong semantic units Kw from surface-level wrong documents Mw in a TriviaQA-based environment with controlled provenance. On Qwen2.5-7B, under No Debate and with Kw=1, increasing Mw from 1 to 8 reduces accuracy from 90% to 47% and raises target-wrong selection from 0% to 48%; holding Mw=6 and varying Kw yields no comparable monotonic trend. Conditional mutual information and Bayesian accumulation provide normative references, while a correlation-neglect model captures residual behavioral weight on copies. Matched Context and round-level states indicate that debate propagates or stabilizes an existing wrong-evidence advantage rather than inherently creating errors. Aggregation- and communication-stage interventions further support repeated exposure as part of the pathway to erroneous consensus. Together, these results support a predictive and falsifiable behavioral-level account of semantic evidence overcounting, without making claims about its internal neural implementation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.