acceptodds
Under review as a conference paper at ICLR 2027

How Many Relation Types Do You Need? Defining a Corpus-Relative Optimum for Text-to-Graph Schemas

Abstract

Every text-to-graph pipeline commits to a schema before extraction, yet prior work offers no principled answer to how large that schema should be. We introduce SchemaGain, which scores a candidate relation inventory by the surprisal that its extracted graph removes from the corpus, minus the language-model cost of describing the schema itself. Across all 19 DBpedia-WebNLG corpora in Text2KGBench, SchemaGain has an interior optimum and aligns with expert-designed inventories (Pearson ; across seeds), while alternative selectors based on clustering, frequency, or direct LLM judgment show weaker agreement and greater variation across seeds. Randomly merging relations through otherwise identical machinery reduces the correlation to . On Text2KGBench's extraction task, using any schema improves F1 by 61% over schema-free extraction and reduces subject hallucination from 41% to below 1%. The cost of larger inventories becomes apparent downstream: in a routing benchmark over the same graphs, relation retrieval peaks near 35 types, and at 60 types a label router's top-1 accuracy is half that at 30. A full sweep takes about four hours on one consumer GPU. Together, these results suggest that schema size is a measurable modeling choice rather than an arbitrary design parameter.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.