ComplexLogicalQA: Benchmarking Knowledge Graph Question Answering Beyond Chain-like Reasoning
Abstract
Knowledge Graph Question Answering (KGQA) benchmarks have substantially advanced the evaluation of multi-hop reasoning. However, our structural audit shows that widely used benchmarks remain heavily skewed toward chain-like queries, with limited coverage of nested logical compositions and no union structures in our sampled analysis. To address this gap, we introduce ComplexLogicalQA (CLQ), a structure-balanced diagnostic benchmark covering nine Existential Positive First-Order Logic (EPFO) query structures. CLQ is constructed through structure-constrained subgraph sampling, followed by LLM-based verbalization and multi-stage verification to preserve executable logical validity while improving linguistic diversity. Experiments across four KGQA paradigms reveal a clear and consistent trend: performance degrades as queries move beyond chain-like projections, with the largest and most stable drops on nested compositions. Union structures are not uniformly hardest under every metric, but they consistently expose weaknesses in complete answer-set recovery. These trends remain visible after controlling for multiple observable non-structural factors. Further diagnostic analyses reveal paradigm-specific bottlenecks in evidence acquisition, graph exploration, and logical-form grounding. CLQ thus serves as a diagnostic testbed for developing more structure-aware KGQA systems.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.