BIRD-Pandas: Diagnosing Cross-Paradigm Divergence in Text-to-SQL and Text-to-Pandas
Abstract
Text-to-SQL is a widely adopted paradigm for natural-language access to structured data, but many real-world analytics tasks involve file-based tables and multi-step transformations that are more naturally handled with code-based approaches. Controlled evaluation of Text-to-Pandas and Text-to-SQL under matched task conditions remains largely absent, making it difficult to isolate paradigm-specific challenges from task content. We introduce BIRD-Pandas, a paired benchmark that derives verified Pandas ground truth from BIRD and enables controlled comparison. Benchmark construction revealed that producing equivalent results across paradigms requires different decision logic due to differences in how the two paradigms allocate execution decisions, a phenomenon we term execution-semantic divergence. Text-to-Pandas trails Text-to-SQL in accuracy at smaller and mid-scale models, and this gap narrows substantially as model scale increases. However, the aggregate gap reflects multiple contributing factors, including execution-semantic divergence and task-specification insufficiency, where natural language questions omit information. Disentangling these factors requires supplying the missing task information and observing the resulting performance change. To this end, we introduce the Logic Completion Framework (LCF), a clarification protocol that measures performance recovery when the model requests and receives such information. LCF closes most of the cross-paradigm gap at intermediate scales. At the largest evaluated scale, SQL and Pandas achieve statistically comparable aggregate accuracy, yet SQL benefits more than Pandas from complete task specification. Analysis of post-LCF failures shows that errors attributable to task-specification insufficiency remain within both paradigms' own error sets, as models tend to overlook consequential ambiguities even when invited to ask clarifying questions, while the net cross-paradigm gap is instead driven mainly by execution-semantic divergence, which is concentrated on the Pandas side. Resources are available at https://anonymous.4open.science/r/Bird_Pandas-3F82/.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.