MedSteer: Clinically constrained representation steering for debiasing medical large language models
Abstract
Large language models are increasingly used for medical decision support and clinical question answering, but their responses can suffer from clinically unjustified demographic bias. Mitigating this bias is complex because attributes like age, gender, and ethnicity may be relevant in certain contexts while introducing inappropriate sensitivity in others. Existing steering methods often fail to separate sensitive attributes from diagnosis-relevant clinical information. To address this, we introduce MedSteer, a clinically constrained, inference-time representation steering framework for medical LLMs. MedSteer estimates attribute-associated representation shifts from diagnosis-conditioned matched pairs and prioritizes steering directions using semantic stability and clinical orthogonality, reducing non-decisive demographic sensitivity while preserving clinical reasoning. It further incorporates an auxiliary lexical control signal from stigma–neutral contrasts without parameter updates. Evaluated across four medical LLMs on sensitivity analysis, long-text entity retention, and medical question answering, MedSteer outperforms an unweighted MeanDiff baseline. It improves average question answering accuracy by 2.6 percentage points across 12 model–dataset combinations, better retains clinical entities, and achieves a superior fairness–utility trade-off. These findings highlight MedSteer as a practical post-deployment solution to mitigate demographic bias in medical LLMs without compromising clinical performance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.