FinAgent-Reliance: Benchmarking Human Reliance and Error Inheritance in Financial LLM Decision Support
Abstract
Existing financial LLM benchmarks typically evaluate whether a model produces the correct answer, but stop before asking whether model failures propagate into human decisions. We introduce FinAgent-Reliance, a source-grounded benchmark and experimental framework for evaluating reliability at the human-AI system level. The framework links authoritative financial evidence to human-validated answers, natural and controlled agent failures, blinded correctness assessment, and separate coding of AI-specific error inheritance. Its frozen Stress-Test v2 contains 47 human-validated records: 12 correct parent responses and 35 isolated controlled failures spanning ten failure categories, with leakage-resistant splits across parent families, companies, and source documents. Stress-Test v2 is difficult for both evaluated open-weight 7B instruction models. Qwen2.5-7B achieves balanced accuracy 0.576 (95% parent-family cluster-bootstrap CI: 0.486-0.681), while Mistral-7B reaches 0.515 on parseable outputs and 0.475 when unparseable outputs are counted as failures. At the failure-typing level, neither model correctly identifies the target category for any of the 20 controlled failures spanning seven semantically subtle categories. Exploratory human studies show that the reliance measurement and evidence-inspection protocols can be implemented in practice; these studies are not used to estimate intervention effects. FinAgent-Reliance evaluates an additional question beyond model correctness: whether a known model error is detected, resisted, corrected, or inherited after it is presented to a human decision-maker.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.