FinDriftBench: A Benchmark for Endogenous Compliance Drift in Financial Multi-Agent Systems Under No External Attack.
Abstract
Large language model (LLM) agents are being rapidly deployed in high-stakes financial decisions such as credit approval, investment, underwriting, and claims adjudication, and with this shift their risk moves from the content level of "saying the wrong thing” to the long-horizon behavioral level of "doing the wrong thing.” Yet almost all existing safety evaluations rely on external jailbreaks, injection, or inducement to trigger violations and are largely single-turn and static, leaving a more fundamental question unanswered: in the absence of any external attack, will an agent, purely by virtue of its own situation, spontaneously deviate from the compliance requirements it was initially given over long, continuous business execution? To address this, we build a multi-agent sandbox that mirrors the functional division of labor of a real financial institution—a system of performance-management, task-execution, and compliance-audit agents—paired with a financial dataset drawn from real business and curated by domain professionals, in order to study how an agent's compliance behavior evolves as it continuously executes financial tasks in the absence of any external attack. We further propose a suite of original metrics tailored to the dynamic execution process, characterizing the onset, frequency, temporal trend, direction, pressure sensitivity, and cognition–behavior separation of compliance-behavior drift, and we attribute the drift to memory forgetting and performance pressure through single-factor ablations. Across 8 mainstream LLMs and 10 categories of financial business, every model exhibits compliance-behavior drift to varying degrees, and we confirm that the violations arise through two pathways: performance pressure and memory forgetting. Moreover, once communication among the agents is opened, the multi-agent system gives rise to team-level collective violations and collusive evasion of the audit. This work provides a reproducible benchmark and diagnostic metrics for the long-horizon compliance monitoring and governance of agents deployed in high-stakes domains.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.