FINCH 2.0: Knowing When Not to Answer in Conversational Financial Text-to-SQL
Abstract
Large language models perform well on single-turn text-to-SQL but degrade on multi-turn, ambiguous, or analytically complex questions. Existing benchmarks test these capabilities separately, and none combines a per-turn decision to answer, ask for clarification, or refuse with compound queries and adversarial variants of the same turns. We address this gap in finance. A seeded planner fixes each conversation's structure over 43 protocols in six series before an LLM writes any text, yielding FINCH 2.0: 16 databases in 7 financial verticals, with 469 tables, 3,674 columns, 13,024 conversations, and 176,860 turns, of which 158,118 are answerable question-SQL pairs. Automated checks validate every turn, and seven domain experts rate 86.0% of 1,287 sampled turns fully correct (95.7% in the test split). On a 2,400-conversation test split, three open-weight models refuse unanswerable questions reliably (recall 89.5–95.3%) but ask on only 34.9–57.7% of turns that need clarification; all fall below the 94.0% Action Accuracy of always answering; on turns a text-only classifier misjudges, Ask recall drops to 20.1–48.8%. SQL quality falls from 0.53–0.57 on easy queries to 0.28–0.33 on hard ones.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.