Beyond Field F1: Concept-Binding Accuracy and the Hindsight Benchmark for Document AI
Abstract
Document AI systems often extract text correctly but assign it to the wrong semantic field, for example, reading “$4,523.67” correctly from a paystub but labeling it net pay instead of gross pay. Value-only matching, which we term Value-Only Recall (Recval), credits a ground-truth value if it appears anywhere in the output, irrespective of its concept key, and therefore misses this failure. We introduce Concept-Binding Accuracy (CBA), a metric family that decomposes joint key-value error into extraction and binding components, and Hindsight, a benchmark of nine datasets across five document domains with hierarchical concept ontologies. For Claude Haiku 4.5 and GPT-4o-mini, Recval overstates joint key-value accuracy by up to 33.3 percentage points: on commercial contracts, GPT-4o-mini reaches 100% Recval but only 66.7% CBA, with 250 misbindings across 50 documents. Even fixed-layout IRS W-2 forms yield 194 (Haiku) and 115 (GPT-4o-mini) misbindings when 44 same-type fields compete, whereas low-ambiguity domains show nearzero gaps. We identify two drivers, concept density and semantic overlap, and confirm both as independent causal effects in a controlled 5 × 4 factorial study on W-2: overlap alone moves binding accuracy by 25 percentage points, and density alone by up to 16 percentage points. An open-weights model (Qwen2-VL-7B) shows the same pattern, with markedly worse binding on the highest-ambiguity domain (CUAD) and comparable binding on low-ambiguity ones. We release all ontologies, scoring code, and experiment notebooks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.