acceptodds
Under review as a conference paper at ICLR 2027

MedCOBE: Diagnosing LLMs’ Collaboration Failures before Clinical Deployment

Abstract

Large language models are increasingly deployed as collaborative partners in the clinical domain, supporting clinicians through interaction. Yet their evaluation often relies on MedQA-style benchmarks that assess LLM knowledge in non-interactive, standalone settings. This evaluation-deployment gap raises a question: does an LLM's standalone performance reflect its ability to support clinicians through interaction? Across 20 LLMs, we observe large performance drops in deployment-like interactions, even on cases that the LLMs can solve in standalone settings (e.g., Claude-Opus-5 exhibits an 11.4% drop in decision accuracy). This shows that standalone evaluation alone is not sufficient to anticipate LLMs' failures in deployment. To address deployment-time failures before deployment, we introduce MedCOBE, a simulation sandbox that evaluates LLMs via deployment-like clinician–AI interactions. MedCOBE quantifies LLMs' interaction behavior: ability to properly react to clinician input in interaction, while isolating failed behavior from knowledge limitations. The MedCOBE score correlates more strongly with collaboration outcomes than standalone accuracy does (rho = 0.820 vs. 0.409). Moreover, MedCOBE produces behavioral profiles that reveal each LLM's interaction tendencies associated with failed collaboration, i.e., frequent sycophancy towards incorrect clinician hypotheses or obstruction of correct ones. We further show that these profiles can enable targeted mitigation of collaboration failure at both model and system levels, establishing MedCOBE as both an evaluation design and an actionable insight for safer deployment in clinical contexts.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.