ImmunoGraphBench: A Large-Scale Context Graph Benchmark for Scientific Reasoning
Abstract
Biomedical knowledge is recorded as conditional claims: a relation between two entities is observed under particular conditions, supported by specific evidence, and associated with particular measurements. A graph interface that exposes only context-free triples can make distinct candidate hypotheses observationally equivalent: they satisfy the same topology, while the observation that distinguishes them is absent. This is an identifiability problem at the reasoning interface, not merely a reduction in text. ImmunoGraphBench makes the problem testable by constructing structurally underdetermined questions where relation-level context breaks the symmetry. We build a biomedical KG of 1.62M triples, 75% of which carry textual evidence or structured metadata, and vary structure and context independently on the same QA instances. Across seven LLMs, adding context to explicit candidate paths raises Medium-level accuracy from 61.35% to 89.56% (+28.20 pp), and the gain persists when the graph is retrieved from the question alone (+14.83 pp). Controls show that the gain depends on relation-matched content, survives evidence-count debiasing, and partially compensates for missing structural support. The results identify relation-level context as the interface signal that restores candidate identifiability beyond topology.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.