Motif-Driven, LLM-Free Benchmark Construction for Question Answering over IT Knowledge Graphs
Abstract
As operational environments become increasingly heterogeneous, organisations are adopting graph-based representations to integrate knowledge across incidents, services, configurations, dependencies and remediation procedures. This creates a need for reliable benchmarks that can evaluate question answering over the resulting operational knowledge graphs. Such evaluation has a fundamental ground-truth problem: the answer is a structured graph object whose correctness depends on the graph's schema and instances, yet benchmark references are commonly produced by language models. This makes it possible for a reference to contain plausible but unsupported relations, making disagreements between a system and its reference difficult to diagnose. We present a graph-first approach to benchmark construction in which ground truth is computed, rather than generated. Starting from structurally informative patterns in a knowledge graph, we compile deterministic graph operations into reference subgraphs and only then generate natural-language questions that describe those references. Language models therefore determine how a question is phrased, but not what constitutes its answer. The same grounding mechanism identifies questions that cannot be answered from the graph and separates schema-level, instance-level, and grounding ambiguities.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.