acceptodds
Under review as a conference paper at ICLR 2027

Benchmarking Financial Agent Safety: Scenarios and Executable Compliance

Abstract

Financial agents must complete useful tasks while respecting permissions and constraints that depend on business state. We introduce an asset-management safety benchmark organized around three deployment settings: customer service, internal assistance, and autonomous production. They differ in requester authority, permitted actions, and requirements for human review. Within thirteen business subscenarios, 877 executable cases tie safety requirements to recorded permissions, business state, and explicit rules. Adversarial and non-adversarial cases test both induced violations and failures during routine requests. Checkers inspect actions, intermediate events, and state to score violations, delivery, and safe completion; selected tasks also test safety under continuation. We evaluate nine models across 144 model-agent-defense configurations. Aggregate scores conceal business-specific weaknesses; agent implementations can improve or worsen a model's safety, as when Hermes lowers Qwen3.6-27B's adversarial violation from 24.6% to 12.8% but raises Qwen3.8-27B's from 7.9% to 19.4%; and first-turn safety can fail under continuation. Added defenses provide different benefits across attack types and can reduce legitimate task completion. A controlled study further separates claimed authorization from legitimate record changes in follow-up turns. Code and cases are available at https://anonymous.4open.science/r/fin-agent-safety.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.