Poisoned Memories: Measuring and Mitigating Demographic Bias Propagating Through Shared Memory in LLM Multi-Agent Systems
Abstract
LLM-based multi-agent systems increasingly rely on shared memory: stores of screening notes, verdicts, and summaries that persist across interactions and are read by every member. We show that this architecture creates a previously under-quantified fairness risk: demographic bias, once written into shared memory, propagates to every agent that reads it. Using a controlled resume-screening testbed with counterfactual name pairs and Qwen3-8B agents, we measure propagation through a committee whose memory is contaminated with group-targeted claims. Starting from a near-fair base model (|DPG|≈0.11), reading contaminated memory raises the demographic parity gap to |DPG|=0.65; five poisoned entries suffice on the probed dose grid, and, on the probed seed, qualified candidates from the targeted group are rejected 3.8× more often. Training-free memory-level governance (provenance gating, individual-only reads) eliminates the group-claim channel on 2 of 3 memory seeds (mean 67% reduction) while improving ranking utility; across 10 memory seeds, however, a name-anchored cascade through individual notes keeps residual bias nonzero on every seed and above baseline on 9 of 10 (mean |DPG|=0.38), its direction following the model's own prior rather than the attacker, and decision-time audit framing amplifies member bias. Adaptive probes bound the defenses: corroborated poison bypasses provenance gating, and individual-level poison defeats all group-claim filters. Memory governance therefore deserves first-class status in multi-agent fairness—a necessary layer, though not by itself sufficient.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.