When a Bad Procedure Outlives the LLM Agents That Wrote It
Abstract
Shared notes let language-model agents build on one another’s work, but they can also carry an invalid practice to agents who never saw where it came from. Understanding collective safety therefore requires separating the emergence of a violation from its propagation and its correction. We study these steps in teams of four LLM analysts performing executable data-analysis tasks under an explicit validity rule. Every published result is independently re-executed, and each team is replaced twice, so successors inherit only agent-written notes. An invalid proce- dure is planted privately with one founder, unexposed teams serve as controls, and paired branches from identical notes either replace them with verified guidance or offer an advisory validity review. Across ten model configurations, violations without exposure were rare in completed runs, and eight configurations adopted the planted procedure at its source. What happened next differed sharply. Among adopting configurations that completed their entire allocation, invalid publications after two complete turnovers ranged from 14% (DeepSeek) to 99% (Qwen), while both GPT-6 configurations never adopted the procedure at all. Replacing the notes with verified guidance produced valid work in all 3,952 complete-block compari- son slots. An accurate review did not guarantee correction: after a failed review, Gemini published a valid result 23 of 27 times, Qwen 0 of 239 times. Initial susceptibility alone does not characterize downstream persistence or correction.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.