FinRiskArena: Time Safe and Multilayer Evaluation of Financial LLM Decision Making
Abstract
Large language models (LLMs) are increasingly considered for high stakes financial decision support, yet conventional benchmark performance provides limited evidence of whether such systems are reliable. Financial risk evaluation must account for temporal validity, differences between imminent and latent risk, operational reliability, evidence grounded decisions, and the validity of the metrics used to assess these properties. We introduce **FinRiskArena**, a time safe and risk sensitive benchmark for multidimensional evaluation of financial LLM decision making. FinRiskArena contains 32,045 company–time episodes from Chinese A share firms constructed under strict point in time information constraints. It separates **imminent distress detection**, which retains legally available near term signals, from **latent distress forecasting**, which removes audit opinions and requires more than 90 days of prediction lead. The benchmark evaluates five complementary layers: **predictive discrimination and calibration, operational reliability, semantic evidence use, counterfactual behavior, and evaluator validity**. A strong structured baseline reaches an AUPRC of 0.439 in the imminent regime, falls to 0.162 after audit opinion information is removed, and reaches 0.047 in the latent regime, showing that superficially similar distress prediction tasks differ substantially in difficulty. LLMs retain meaningful risk ranking signal, but predictive performance does not consistently align with structured output reliability or semantic auditability. Counterfactual tests reveal heterogeneous behavior across perturbations, while a separately frozen blinded financial expert validation shows that selected automatic counterfactual scores diverge from expert semantic judgments. FinRiskArena therefore treats trustworthy financial LLM evaluation as a multilayer measurement problem in which predictive quality, operational compliance, semantic reliability, counterfactual behavior, and evaluator validity must be assessed separately.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.