acceptodds
Under review as a conference paper at ICLR 2027

Aligned Alone, Misaligned Together: Forecasting Adversarial Capture in LLM Agent Populations

Abstract

AI safety evaluations typically assess models one at a time. But language-model agents increasingly operate in populations, where they read and respond to one another's decisions. A model that makes reliable decisions on its own may behave differently when surrounded by other agents. Can individually reliable agents be collectively manipulated, and can we predict how strongly the population will respond? We study this question in a security-triage task, where populations of language-model monitors decide whether to escalate or dismiss alerts. We introduce a committed minority of agents that consistently advocates for one decision and measure how its presence shifts the population's behavior. We find that alerts judged almost identically by an isolated agent can produce sharply different collective outcomes. Individual audits therefore cannot reliably predict population behavior. Yet collective vulnerability is not unpredictable: using only benign, adversary-free behavior, we calibrate a response function that forecasts how strongly a population will shift under a future attack, before the attack is run. We next examine how interaction and time shape these outcomes. Allowing agents to observe one another's reasoning can weaken adversarial influence, but does not reliably prevent capture. However, capture need not be permanent. After adversaries are removed or replaced, mean population behavior returns toward its original level, even after five or thirty rounds of pressure. These findings are consistent with a calibrated single-stable-state model, in which adversarial pressure shifts the operating point without creating a second self-sustaining state. The same calibration predicts attack responses in populations of 24 and 48 agents. Alignment in isolation does not guarantee alignment in a population. But a population's behavior before an attack can reveal how it will respond when one arrives.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.