LLM Extractor Errors Are Not Random: Protagonist Bias in Knowledge Graph Extraction
Abstract
Large language models have become the prevalent tool for building knowledge graphs from scientific literature. It turns unstructured text into entity-relation graphs which enhance literature-based discovery, biomedical question answering, and drug repurposing. However, the resulting graphs remain far from complete and accurate, and common remedies such as sampling and voting implicitly treat extraction errors as independent noise. Through cross-corpus error analyses, counterfactual experiments, and mechanistic analyses, we show that these errors are structured: relations between a document's supporting entities are extracted systematically worse than those involving its most-mentioned entity, its protagonist. We trace this to document-level salience: an entity's prominence across the whole document, not just in the sentence that states the relation, shapes whether the model opens a relation with it, and also inflates the model's confidence, so frequency- or confidence-based selection reinforces the bias. We therefore separate proposing relations from verifying them: an evidence-anchored extractor re-reads each document sentence by sentence to propose candidates, and a selector keeps only those supported by local evidence. On BioRED, this improves Qwen3.5-27B by 0.07 F1 over standard fine-tuning and raises supporting-entity precision from 0.54 to 0.68 without loss of recall.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.