acceptodds
Under review as a conference paper at ICLR 2027

GAMUT: A Benchmark for Factual Completeness with Two-Level Meta-Rubrics

Abstract

We introduce GAMUT (Grounded Assessment of Multimodal Factuality), an everyday deep research benchmark for factual completeness, built from naturally phrased questions concerning everyday subjects, yet requiring multi-step research and multi-paragraph answers. Evaluating these answers exposes a fundamental tension in rubric-based evaluation. A faithful rubric must express the structure of the space of good answers: open-ended sets of acceptable options, ordered processes, and relative importance. Consistent automatic grading, however, favors narrow, independently checkable criteria. GAMUT addresses this tension with a two-level meta-rubric design: a structured meta-rubric captures answer requirements at authoring time, which are compiled by fixed mechanical rules into a flat binary checklist for evaluation. The benchmark comprises 1,813 questions grounded in real-world images across 10 domains, each paired with an evidence-backed rubric verified by expert human annotators. Evaluating 14 frontier and open-weight models, we find GAMUT challenging and highly discriminative, with scores ranging from 5.1% to 58.7%, and broadly consistent rankings across three LLM judges.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.