Right Answers, Poisoned Provenance: What Accuracy Misses in LLM-Based Multi-Agent Reasoning
Abstract
Large language model (LLM) agents increasingly answer questions as a group, exchanging answers, rationales, and cited evidence that no single agent holds in full. Task outcomes and answer persistence do not establish the reliability of the cited evidence that humans and downstream agents may reuse. We study a distinct failure, provenance poisoning: adversarially planted misinformation enters honest agents' cited trails, whether or not the group answers correctly. We define a judge-free provenance-poisoning rate (PP) computed from debate logs and known planted-source identities. One adoption can poison the trail without changing the plurality answer. Across two multi-hop QA corpora and two open-model families using original source IDs and audited misinformation, one contaminated agent moves answer-and-evidence (Joint) F1 by at most .032 in absolute terms, yet its misinformation enters honest agents' citations in 12–50% of instances. The Joint-F1 change includes a 19.4% relative decline in one setting. Correct answers can also retain poisoned citations. On MINT-CWQ with Qwen3-8B, poisoning reaches 64.3% with two contaminated agents and opaque source IDs. In a separate original-ID reuse experiment, a new reader answers from passages cited by agents matching the group output, including groups that return no answer. Across two reader families, removing a planted passage adopted from a peer significantly improves Answer F1 by 4.3 points for Qwen and 2.0 for Mistral over all 250 MuSiQue questions, and helps significantly more than removing a length-matched legitimate passage. Poisoning rates vary across models in ways answer scores cannot distinguish, and a defense that checks peer claims has model-dependent effects. Multi-agent safety evaluation should measure the provenance of cited evidence, not only the answer.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.