Correct Scores, Unsupported Conclusions: Benchmarking Evaluation under Incomplete Semantic Reference
Abstract
An evaluation score can be computed exactly against the available reference and still fail to support the scientific conclusion drawn from it. We formulate model comparison under incomplete semantic reference as an evidence-qualified evaluation problem: when observational proxies, current snapshots, or incomplete records substitute for the required semantic reference, what evidence is sufficient to certify which model is better? Holding predictions fixed across structured domains reveals that proxy distortion alone does not determine conclusion validity: selected-only reference drops chess forest set from to while leaving the model winner unchanged, yet the identical substitution reverses model selection in Planning Domain Definition Language (PDDL) tasks. This divergence establishes that reference fidelity and conclusion fidelity are distinct: exact reference identification is sufficient but not necessary if every compatible reference induces the same model comparison. We introduce a typed benchmark across public and executable sources spanning stable, reversed, recoverable, certifiable, and uncertifiable regimes, and operationalise an executable certification protocol that determines whether available evidence already certifies the comparison, target-matched recovery suffices, additional evidence must be acquired, or the comparison must remain \Unknown. In process panels, target-matched recovery resolves ranking conclusions despite imperfect reference reconstruction; in controlled and natural-history panels, exact query certificates determine whether evidence certifies the query, requires witnesses or exhaustion, or yields \Unknown. Holding outputs fixed, stronger references can preserve rankings, alter pairwise comparisons, or reverse top-system selection, with the latter observed in both formal planning and a contemporary speech-recognition panel evaluated against independently corrected transcripts. Evaluators should therefore determine whether available evidence is sufficient to certify the declared model comparison, rather than forcing an ungrounded point reference and mistaking predictive confidence for semantic warrant.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.