STAR-Bench: Evaluating Anti-Money Laundering Agents for Regulatory Reporting Workflows
Abstract
Anti-money laundering (AML) investigations are regulated tool-using workflows that end in a suspicious transaction report (STR), and they are harder than general tool use in four ways: tools reading the same fields serve different regulatory roles, reporting is an obligation, evidence must carry across turns, and models are chosen by proxies whose bearing on this work is unknown. Existing function-calling and financial benchmarks do not score whether an agent carries out a mandated reporting step. We introduce STAR-Bench, a compliance-oriented benchmark derived from an AML agent platform and expert-verified scenarios: 23 AML tools and 1,308 evaluation cases spanning single-turn tool use and multi-turn STR workflows, where every tool executes against the database. Experiments on 28 open-weight configurations show that single-turn tool accuracy does not resolve workflow readiness: configurations within seven points on single-turn tool hit differ by thirty-two points in STR workflow completion, and the best completer ranks 17th of 28 on single-turn hit. Workflow failures concentrate at the STR-writing turn and persist under an oracle that injects executed tool outputs, while end-to-end execution costs parameter grounding rather than tool selection. Reporting tools are about as hard to trigger as analysis tools on average but fail differently: three quarters of STR-field validation failures are calls never made. General function-calling rank, finance specialization and schema language give no reliable ordering of these results, and reasoning mode lowers them on average under our output budget. We release the STAR-Bench dataset, executable evaluation environment, evaluation code, and AML agent platform.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.