acceptodds
Under review as a conference paper at ICLR 2027

CrowdRAG: Source-Density Defense Against Sybil Poisoning in Multi-User Agent Memory

Abstract

LLM agents increasingly operate over shared memory pools where multiple users both contribute and retrieve, a setting that lets a single user's contribution become evidence that any member's agent retrieves and acts on. This shared write access enables Sybil poisoning: three fake accounts, 1% of users and about 10% of a typical topic's contributors, already fill 53% of the top-5 retrieved passages. Existing content-level defenses systematically fail because honest paraphrases and adversarial claims about the same topic are structurally indistinguishable in embedding space. We propose CrowdRAG, a two-layer defense that exploits the orthogonal signal multi-user settings provide: the number of distinct users independently supporting each claim. CrowdRAG first filters claims that lack support from enough distinct users, then resolves remaining contradictions through a provenance-aware knowledge graph. With three Sybil accounts and GPT-5.4, CrowdRAG reduces attack success from 63.6% to 1.3% at 96.1% accuracy; attack success stays at or below 3.9% across three LLMs and, with a label-free similarity calibration, across four retrievers. We also characterize a density cliff: on topics with no more honest contributors than Sybil accounts, no support threshold can separate the two. Our results show that user provenance is a practical trust signal complementary to content-level defenses.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.