acceptodds
Under review as a conference paper at ICLR 2027

CAST: Community-Adaptive Social Red-Teaming of LLM-Agent Societies via Population-Level Preference Optimization

Abstract

Large Language Model (LLM) agents are increasingly deployed as persistent participants in shared social environments, where they continuously interact with each other for content generation and exchange. Such an ecosystem has given rise to population-level vulnerabilities, where harmful or inappropriate content can affect multiple agents instead of a single model in isolation. To study this phenomenon, existing work either adopts a predefined, heuristic claim to trigger the threat, or generates and optimizes adversarial inputs against a standalone model. They all ignore the impact of the LLM-agent community and agentic social behaviors on misinformation propagation. To address this gap, we introduce CAST, a community-adaptive social red-teaming framework. The key idea of CAST is to incorporate the feedback from the target community as the supervision signal to optimize the adversarial input, which can accurately and comprehensively reflect the population-level vulnerabilities in the LLM-agent society. Specifically, given a sampled community context, CAST generates diverse naturalistic adversarial posts and converts them into pairwise preferences by measuring their effectiveness from agents' behavioral responses. The adversarial input is further optimized with Direct Preference Optimization according to the population feedback. Our evaluation shows that CAST significantly improves the red-teaming performance compared to baselines. We further find that agent-level behavioral traits substantially alter susceptibility under identical attacks and exposure. Together, these results prove that population-level vulnerabilities in LLM-agent societies are jointly shaped by adversarial content, community context, and agents' characteristics.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.