Literature-Grounded Language Models for Auditing Biomedical Ontologies
Abstract
Biomedical ontologies encode clinical distinctions that downstream systems inherit, but auditing them is difficult: many apparent defects are legitimate modeling choices, while many real defects require both graph-level and domain-level evidence to recognize. We study whether language models can act as triage systems for biomedical ontology audit when their prompts expose progressively richer context: concept names alone, release-specific ontology neighborhoods, and ontology neighborhoods grounded in recent clinical literature. Our setting is deliberately narrower than automated ontology repair. The model's task is to propose candidate defects; validity, actionability, and novelty must be established by independent review. We build an evaluation framework around SNOMED CT clinical branches and UMLS source-preservation checks, with frozen request packets, repeated model runs, exact prompt and response logging, failure accounting, blinded proposal export, and reviewer-gated analysis. A retrospective source comparison across twelve clinical branches identifies 21 concepts whose measured SNOMED CT US records changed between March and September 2026, including 15 with changed transitive ancestry, providing naturally occurring release-change material for audit evaluation. In sleep-apnea development experiments, open-ended model audits surfaced candidate terminology issues, including a baclofen mechanism discrepancy later supported by external sleep-medicine expert correspondence: baclofen is pharmacologically a GABA-B agonist, whereas the examined terminology context represented it with a GABA-A receptor agonist disposition. These findings motivate, but do not replace, blinded adjudication. The full study evaluates four frontier models across three context conditions, repeated runs, and multiple clinical branches. We report proposal yield, precision under independent human labels, run-to-run variability, model-specific effects, failure rates, and cost. By separating candidate generation from validation, and by treating ontology source records, clinical literature, expert feedback, and curator disposition as distinct evidence layers, the framework tests a practical question for high-stakes biomedical AI: can language models improve the prioritization of terminology defects without collapsing plausible explanation into proof?
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.