Benchmarking Interactive Clinical Diagnosis under Patient Stigma and Selective Information Disclosure
Abstract
The focus of medical artificial intelligence (AI) is shifting from static knowledge reasoning to interactive clinical diagnosis. However, existing frameworks assume unconditionally cooperative patients who transparently surrender clinical facts, contradicting real-world consultations where disease stigma, privacy concerns, and fragile interpersonal trust induce selective information disclosure. We propose , the first benchmark that formalizes interactive diagnosis as an imperfect-information, general-sum game to evaluate medical AI under disease-stigma awareness. By mining public medical corpora, we first normalize disease stigma by 185 canonical Clinical Stigma Units (CSUs) with scaled relative sensitivity. Then, we construct a neuro-symbolic agentic patient that mandatorily invokes a dedicated tool before responding to determine its disclosure posture towards doctor-queried information. This tool dynamically tracks the patient’s trust in the doctor via an evolving Beta belief distribution and stochastically samples disclosure strategies based on the posterior tail probability of trust overcoming stigma sensitivity, thereby mirroring human bounded rationality and emotional hesitation over rigid cognitive principles. Evaluations across 13 frontier LLMs reveal that ideal-patient evaluations overestimate the clinical robustness of current AI, evidenced by degraded inquiry efficiency, trust-building, and diagnostic accuracy under selective-disclosure postures. Dedicated agentic scaffolding improves sensitive-evidence elicitation but remains capped in diagnostic accuracy, leaving room for trust-aware, reliable clinical AI.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.