Ask or Guess? A Minimal-Pair Benchmark for Clarification-Seeking in LLMs
Abstract
Deployed assistants receive requests that are underspecified, ambiguous, premised on false facts, contradictory, or unknowable. For many of these, the reliable re- sponse is not a refusal but a question. Abstention benchmarks measure whether models decline to answer; interactive benchmarks measure whether they can clar- ify across multi-turn dialogue. Harder to compare is the single-turn decision boundary itself — whether a model answers, asks, or declines on a defective re- quest, whether it over-clarifies on requests that were answerable to begin with, and what even counts as a correct response on each side. We introduce CLARPAIR, a benchmark of 162 paired requests spanning five defect types: each defective request is paired with a twin that differs only by the missing piece, so answer- ing, clarifying, and abstaining can be scored jointly without confounding task difficulty. Pairs are LLM-drafted, verifier-filtered, and re-verified by a stricter, different-family verifier whose answerability test flags 40 twins that silently re- quire live information — a slice we score separately rather than mislabel as over- clarification. We score responses under an explicit behavior contract — naming the defect, asking back, correcting a premise, and conditional help all count as ap- propriate handling; only committing to an unsupported specific answer counts as a guess. Evaluating nine model configurations across five families and two access regimes, we find defective-request handling at 87–94% and correct answers on 90– 97% of verified-answerable twins, yet joint pair success — both sides right on the same item — reaches only 80–90%. Failures concentrate where the contract bites: on contradictory requests models silently resolve the conflict instead of surfacing it (as low as 17% appropriate), and on requests that only tools could answer, text-only API models commit unsupported answers at 39–56% (62–68% for tool-equipped agents, some plausibly legitimate retrievals). Twin-side errors decompose into rare over-clarification (2–8%) and a newly visible class of wrong answers (0–8%). Un- der a coarser single-label rubric the same responses looked substantially worse, which is itself a finding: the scoring contract materially changes measured conclu- sions. We release the benchmark, the three-stage verification pipeline, all behavior labels, and the scoring harness.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.