FairWatch: Interaction Topology as a First-Order Variable Shaping Safety and Fairness in Multi-Agent Deliberation
Abstract
AI safety practice assumes that aligning individual models ensures safe collective systems. We show this compositionality assumption fails in deliberative agent networks. System-level safety and fairness depend directly on interaction topology: the communication graph, execution order, and aggregation rule coordinating agent behavior. We audit synthetic credit-underwriting decisions across six open-weight models spanning 3B to 72B parameters, holding prompts, profiles, and weights fixed while systematically varying topology across four standard archetypes. Three structural failure modes emerge. First, in sequential chains, agent execution order generates severe decision variance on ambiguous borderline files, producing outcome reversals for nearly one in five affected applicants, concentrated in a small fraction of boundary contexts. Second, downstream agents abandon private domain signals in nearly all conflicting cases across both model families (Meta LLaMA and Alibaba Qwen 2.5), collapsing epistemic diversity into herd consensus regardless of model scale. Third, under hierarchical judge synthesis, a large-scale model exhibits an adverse impact shortfall for protected applicants that does not survive multiplicity correction, while a matched majority-vote readout remains compliant on the same profiles. These findings show that component-level alignment does not guarantee system-level safety, establishing interaction topology as an essential first-order audit variable for multi-agent governance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.