Know Thyself: Concept-Level Diagnosis for Financial Large Language Models
Abstract
Large language models (LLMs) are increasingly used in finance, but existing benchmarks provide limited evidence of what financial capabilities a model has actually mastered. As a result, models with similar benchmark scores may differ sharply in their underlying strengths and weaknesses. We introduce FinCDM, a framework for diagnosis at the concept level that combines the official CPA and CFA concept taxonomies, mappings between questions and concepts verified by experts, model response patterns, and a replaceable capability estimator.}We construct a bilingual benchmark suite derived from the CPA and CFA examinations, consisting of CPA-KQA (Chinese, 1,260 questions, 70 concepts) and CFA-KQA (English, 738 questions, 41 concepts). Across 20 representative LLMs, we find that score-equivalent models are often not capability-equivalent, and that benchmark imbalance and linguistic limitations can further distort apparent competence. We further propose SkillTune, a diagnosis-guided data selection strategy for fine-tuning that targets each model's weak concepts. SkillTune improves multiple LLMs and, in a controlled comparison, outperforms alternative data selection strategies. These results support financial LLM evaluation based on concept-level profiles and diagnosis-guided improvement. Datasets and scripts will be publicly available.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.