When Verbose Responses Mislead Evaluators: Analytic Debiasing for Generative KGQA
Abstract
Knowledge Graph Question Answering (KGQA) aims to answer natural-language questions over a knowledge graph. Classical KGQA systems typically return a single canonical answer entity or an executable logical form, so exact-match evaluation is usually well aligned with their outputs. Recent systems, however, increasingly generate free-form natural-language responses. This makes string matching less reliable: it can reject correct answers with mismatched surface forms and accept incorrect responses that merely mention answer-like entities. We formalize this issue as evaluator bias in generative KGQA. In our framework, observed scores reflect not only latent semantic correctness, but also false-negative mismatch and false-positive contamination. We then propose the Analytic Debiasing Evaluator (ADE), a simple, interpretable, and judge-free framework. ADE starts from matches accepted by a string-based evaluator and assigns each accepted match a reliability score based on its risk of lexical contamination. ADE estimates lexical-contamination risk using matched negative controls and discounts accepted matches that are likely to be contamination-prone. Experiments on human-annotated meta-evaluation sets show that ADE aligns more closely with human judgments than string-matching evaluators. ADE also achieves competitive agreement with LLM-as-Judge evaluators while being deterministic, faster, and free of API cost. In leaderboard analyses, ADE produces contamination-aware rankings that better match human semantic judgments, suggesting that it can help diagnose when biased evaluators distort model rankings.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.