acceptodds
Under review as a conference paper at ICLR 2027

Rethinking Fairness in Multimodal Large Language Models for Medical Imaging

Abstract

Multimodal large language models (MLLMs) now answer clinical questions and draft radiology reports at expert-competitive quality, yet early audits show that their outputs differ in quality across demographic groups. Fairness in medical imaging has been studied extensively, but almost entirely for models that predict a label from an image alone, and the few studies of medical MLLMs are purely empirical and narrow in scope. As a result, there is no theoretical account of how demographic bias arises in an MLLM, which is conditioned on both an image and a textual query and produces open-ended text, and no systematic benchmark to measure it. We address both gaps with MedFairVLM. We first define a fairness criterion for open-ended generation as the disparity in LLM-as-a-judge scores between demographic groups. Building on it, we derive an upper bound on the fairness gap of an MLLM that decomposes it into five sources of bias: disparities in the image distribution, in the query distribution, in the two conditional dependencies between image and query, and in text generation given both inputs. We extend the bound to distribution shift and provide an estimator that quantifies each source from data. We then benchmark ten open-source and commercial MLLMs on eight datasets spanning four imaging modalities, four organs, and three sensitive attributes, across closed-ended visual question answering (VQA), open-ended VQA, and report generation, under zero-shot, fine-tuning, and distribution-shift scenarios. Bias is prevalent across all models and settings; under fine-tuning, fairness behavior is highly correlated with the bias sources identified by our theory; existing fairness methods fail to mitigate bias consistently; and distribution shift makes fairness harder still to control.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.