acceptodds
Under review as a conference paper at ICLR 2027

RegBench: Source-Grounded Benchmarks for Regulatory Cross-Reference Reasoning

Abstract

Compliance work in regulated domains requires tracking cross-references across sections and applying the resulting chain to numeric or categorical decisions. Existing benchmarks test long-document retrieval and multi-hop QA, but the benchmarks reviewed here do not isolate traversal over explicit regulatory cross-reference graphs. We introduce RegBench, a specialised benchmark comprising **827 expert-level questions grounded in 4,766 atomic-fact propositions** from DNV Ship Rules and the U.S. Basel III capital regulation, 12 CFR Part 217. Frontier systems reach only **17–49%** strict accuracy on DNV and **17–57%** on Basel, with clear tier degradation as required context scope increases, and an independent judge from a third vendor reproduces the ranking. Three published graph-RAG methods further collapse to ≤17% strict at pilot scale. To scale such evaluations, we present a novel framework that converts regulatory cross-reference graphs into source-grounded QA items with minimal human effort. This is achieved through a graph-anchored generation pipeline that samples reference chains to synthesise scenarios, and a calibrated audit (v5d) that bounds quality without expert review of every question. Our automated audit agrees with full-set SME annotation on **99.0%** of audited items and **99.1%** of atomic facts on DNV, and transfers to 12 CFR Part 217 without tuning. Together with independently validated answer grading, these results support strict atomic-fact evaluation at benchmark scale. Code and data are available at https://github.com/regbench/regbench.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.