EverydayMedQA: A Consumer-Oriented Multimodal Benchmark for Medical Question Answering in Everyday Settings
Abstract
Conversational AI assistants are increasingly used for healthcare guidance, both before and after consulting a clinician, to prepare for appointments, clarify medical advice, and seek additional information. Recent large multimodal models (LMMs), including omni-modal systems that process text, image, audio, and video, can now respond to medical queries expressed through many forms of input. Yet most medical AI benchmarks remain clinician-oriented: they use expert terminology, rely on clinical-grade inputs, cover narrow medical domains, and typically evaluate a single modality. As a result, how these models handle the informal, consumer-generated queries common in real-world telemedicine remains poorly understood. We introduce EverydayMedQA, a consumer-oriented, multi-modal medical question-answering (QA) benchmark of 28,025 question-answer pairs across six input types: text, image, audio, speech, video, and document. Every question is phrased as a layperson query and paired with consumer-grade media rather than clinical instrumentation, spanning a broad range of medical conditions, specialties, anatomical regions, geographic contexts, and consumer question intents. Certified physicians contributed to benchmark construction through data provision and review of cases requiring clinical judgment. Each item is available in both multiple-choice and open-ended formats, evaluated using a reference-based protocol. Evaluating a comprehensive set of 23 leading models spanning omni-modal models, LLM, VLM, medical VLM, and audio LLMs, we observe that medical question answering performance does not transfer uniformly across input modalities. An omni-modal model that performs well on text may still struggle with patient questions presented in speech, images, audio, video, or documents. We also find that multiple-choice and open-ended questions probe different abilities: a model may select the right options but still give an incomplete free-form medical answer for the same sample. In many cases, models give fluent and relevant-sounding answers that remain clinically incomplete or incorrect. These results suggest that clinician-curated, single-modality benchmarks may overestimate model readiness for consumer-facing telemedicine. To our knowledge, EverydayMedQA is the first consumer-oriented medical QA benchmark to span these six input types under a unified evaluation protocol. Our benchmark and evaluation script will be made publicly available.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.