FinVQA-Chart: A Diagnostic Benchmark Isolating the Multimodal Fusion Bottleneck in Financial Chart Understanding
Abstract
Vision–Language Models (VLMs) report high aggregate accuracy on multimodal benchmarks, yet a single accuracy number cannot localise where a model fails in visual perception, in language reasoning, or in the cross-modal integration that binds them. We introduce FinVQA-Chart, a diagnostic benchmark that separates these three abilities via modality-controlled question triplets: for each underlying problem, we construct a visual-only, a text-only, and a fusion variant, where the fusion variant is solvable only by combining a cue present exclusively in the image with a cue present exclusively in the text. The benchmark contains charts rendered from authentic OHLCV data for U.S. equities (, seven volatility regimes) paired with five-option multiple-choice questions stratified by reasoning category, modality, and difficulty. We define Fusion Efficiency (FE) as multimodal accuracy normalised by the stronger unimodal channel, and, unlike prior "diagnostic" benchmarks, we provide a construct-validity study showing that FE carries signal beyond raw multimodal accuracy and beyond question difficulty: on difficulty-matched items, human Fusion Efficiency reaches while the strongest model stays at , showing the gap reflects integration, not item difficulty. Evaluating contemporary VLMs (frontier, reasoning-specialised, chart-specialised), the strongest model achieves overall, points below human experts, with FE vs for human experts (chance-corrected). Neither extended chain-of-thought nor chart-domain specialisation closes the gap; error stratification localises the residual to numerically grounded operations rather than perception. We release the dataset, an identity-controlled contamination-safe variant, a real-chart transfer set, the evaluation harness, prompts, and baselines. All artefacts derive from authentic market data via standardised rendering with template-controlled language; we provide evidence that the design does not create trivial shortcuts (a model fine-tuned on in-template data reaches only , below frontier zero-shot) and that findings transfer to real charts (rank correlation r=).
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.