SemBench-Hard: Evaluating Query Translation and Operator Grounding in LLM-Backed Semantic Query Processing
Abstract
Structured LLM workflows must turn a request into steps, and each intermediate step must return not only a decision but a grounded artifact that later steps can check and consume. Semantic query-processing engines (SQPEs) expose this need: they answer analytics questions over text-rich records with LLM-backed FILTER, JOIN, MAP, and RANK operators. An NL-facing SQPE can fail in two places: when it compiles a question into an operator program, and when an operator returns a decision without the grounding, such as evidence sentences or attributes, that its output contract requires. SemBench-Hard makes both harder and evaluates them separately. Its translation axis withholds the step-by-step plan from the question and scores generated programs by their final output; its execution axis runs reference programs through three system adapters with three backbone LLMs on paired Base and Hard operator tasks over four public datasets, where Hard adds distractors that make the right grounding harder to find, explicit constraints, output schemas, and deterministic grounding checks. On the translation axis, many accepted programs differ from the reference, so exact-match program scoring is too strict, and goal-only queries lead two of four translators to treat MAP’s string output as a structured object. On the execution axis, Hard tasks lower strict (reference-exact) acceptance from 86.0% to 60.1% (−25.9 points); crediting grounding-only rejections at the rate human annotators judge broadly valid leaves −16.1, and accepting every grounding-only rejection leaves −6.3. The main finding is that operators mostly keep the decision but often return grounding that fails the contract’s check; annotators judge about a third of these rejected outputs insufficiently grounded or wrong, three times the rate among accepted outputs. The loss persists: chains (LOTUS × Qwen3) stay 25–31 points lower under Hard at every length, and bounded task-local tools add 1.0 point at 2.24 times the API cost (4.0 with a SWE-bench helper). SQPEs should therefore schema-check generated programs and check each operator’s grounding.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.