PRISM: Pareto Routing over Input Structure and Models for Budgeted Table Question Answering
Abstract
Table question answering (Table QA) requires a model to derive an answer from a variably-structured table and a natural-language question. Although general-purpose LLMs, including GPT-class systems, and smaller language models (SLMs) can solve many such questions, a fixed model and fixed serialization force a deployment to choose between accuracy, cost, and latency before the difficulty of each input is known. Thus, we introduce PRISM, Pareto Routing over Input Structure and Models, which treats the model–representation pair as the primary decision and uses a second, evidence-gated call only when the first response fails a prespecified acceptance screen. The screen combines answer availability, evidence consistency, and reported confidence; a deterministic quota bounds mean API use by 1.25 calls per question. We derive a risk-priced continuation boundary, conditional representation–model crossover guarantees, and a risk bound that accounts for denied recovery. Completed experiments reach 83.02% strict accuracy on HiTab and 78.04% denotation accuracy on WikiTableQuestions (WTQ). On 851 paired WTQ questions, PRISM yields 43.43% more correct answers per accounted dollar than Pro/CSV, while trading away 8.96 accuracy points and adding 2.99 seconds of logical latency. The results support selective recovery and make its resource trade-offs explicit; they do not claim calibrated uncertainty or universal Pareto dominance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.