FCMBench: A Large-scale Financial Credit Multimodal Benchmark for Real-world Applications
Abstract
We introduce FCMBench, a large-scale multimodal benchmark designed for understanding identity-linked documents in financial credit evidence review applications. It contains 4,922 real-captured images and 11,045 VQA instances covering 26 certificate types, three perception tasks, four credit-specific reasoning tasks, and ten real-world robustness challenges. To address the privacy–realism trade-off in financial data construction, we develop a synthetic-to-physical pipeline that creates linked fictional applicant profiles, generates certificate templates, physically fabricates documents, and recaptures them under controlled scenarios. Task design is grounded in interviews with senior credit reviewers, and a two-stage audit refines data quality. We evaluate 39 proprietary and open-source vision–language models from 14 organizations. Results show substantial performance gaps among current VLMs, with overall F1 scores ranging from 24.60 to 71.57 (mean 50.9; SD 12.6 across models). Matched robustness evaluations further reveal consistent degradation under challenging conditions, particularly for multi-document compositions, small regions of interest, and defocus. Extensive experimental results demonstrate that FCMBench is challenging and highly discriminative for identifying performance gaps among vision–language models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.