Bayesian Agentic Medical Diagnosis
Abstract
Efficient sequential clinical diagnosis typically requires maintaining calibrated beliefs over diseases and selecting tests that maximally resolve uncertainty. Current LLM-based diagnostic agents handle both tasks inside the model itself, through accumulated context, verbalized confidence, or prompt sampling, inheriting the LLM's miscalibration and discarding the principled structure of Bayesian decision theory. We revisit this architectural miscalibration and introduce Bayesian Agentic Medical Diagnosis, abbreviated as “BAMD", which confines the LLM to a predictive role while delegating belief maintenance and test selection to an external Bayesian engine. Beliefs are updated via Bayes' rule after each observation, and test selection is governed by information-theoretic operators on the evolving belief state. On four clinical benchmarks, BAMD integrated on an 8B Bayesianized Llama model running on a single local GPU matches or exceeds prompt sampling with a locally deployed 120B LLM across all four datasets, while producing the best-calibrated probability estimates of any LLM-based approach we compare against. These results highlight that the limiting factor in LLM-based sequential diagnosis is architectural, not scale-related, and that a small model embedded in the proper framework can outperform much larger proprietary systems accessed through commercial APIs.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.