FinRiskAtlas: A Decision-Aligned Benchmark for Financial Risk Review and Evidence Seeking
Abstract
Financial risk review requires more than knowledge of financial rules: a model must produce supported outputs from available evidence and decide when to seek additional information. We introduce FinRiskAtlas, an expert-reviewed Chinese-language benchmark for evaluating these complementary capabilities. Its static component contains 9,742 instances across 42 knowledge families and eleven review operations. Each operation is specified before instance construction by its visible evidence, decision object, required output, and scoring rule. FinRisk-Ask complements these tasks with offline evaluation of 680 pre-action states reconstructed from 104 professional review trajectories. Models receive only the record and history available before the reviewer acted, choose Ask or Proceed, and generate one focused evidence request when asking. Evaluation separates agreement with recorded reviewer actions from request alignment with expert-verified evidence needs. Later observations support target verification but remain hidden during inference. Across 33 model configurations, the knowledge leader attains the highest reported score on only one of the eleven review operations; selecting it for information extraction leaves an 18.01-point gap to the operation leader. On FinRisk-Ask, well-targeted requests can coexist with weak agreement on recorded decisions to proceed. FinRiskAtlas provides operation-specific profiles and evidence-seeking diagnostics for comparing models across review tasks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.