Reasoning or Retrieval? Diagnosing LLM Failures on Temporal Knowledge Graphs with Ground-Truth Evidence
Abstract
When an LLM fails at multi-hop reasoning over an event history, is the model unable to reason or did retrieval fail to provide the evidence it needed? Existing benchmarks cannot disentangle these explanations because the events that caused an outcome are unknown. We introduce a controllable synthetic temporal knowledge graph (TKG) generator that constructs event histories from explicit multi-hop temporal patterns, calibrates them to real event statistics, and records the ground- truth evidence behind every event. This lets us directly separate retrieval from reasoning failures and measure how well retrieval reaches the evidence a query requires. Benchmarking seven LLMs against graph- and rule-based forecasters under a range of retrieval strategies, we find that strong LLMs can overcome increasing hop length when given the right context (the pattern’s ground-truth evidence together with one solved example) while neither suffices alone. For example, on 3-hop patterns, GPT-5.6-terra reaches 3.7 the Hits@1 of the GNN forecaster RE-Net. The dominant obstacle is instead retrieval. Under realistic retrieval methods, 82–98% of errors on 2- and 3-hop patterns occur because necessary evidence never reaches the model, and the best existing and proposed retrievers close at most 39% of the gap to the right context. Models remain vulnerable to frequency and recency shortcuts even when the correct evidence is present. Together, these results recast multi-hop temporal reasoning as primarily a retrieval problem. Strong LLMs can often perform the reasoning, but realizing that capability requires retrieval methods that reliably surface the evidence that matters.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.