PEMA: EXTERNAL ADAPTATION OF FROZEN 3D MEDICAL MLLMS
Abstract
Pretrained multimodal large language models (MLLMs) perform poorly on specialized medical question-answering tasks, while task-specific fine-tuning can substantially improve accuracy at high computational cost. We investigate whether similar performance gains can be obtained by training a small adapter without updating the pretrained MLLM. We introduce Performance-Enhancing Multimodal Adapter for MLLMs (PEMA-MLLM, abbreviated PEMA), an adapter placed between the multimodal projector and the language model. PEMA modifies the projected visual tokens conditioned on the question, while the vision encoder, multimodal projector, and language model remain frozen. We validate our proposed concept on M3D-LaMed models with Phi-3-4B and LLaMA-2-7B backbones on closed-ended 3D-RAD Tasks 4–6. On Phi-3, PEMA achieves 67.40% mean subtask-macro accuracy, compared with 29.86% for the frozen model and 68.79% for the released fine-tuned M3D-RAD checkpoint, while training only 12.08M parameters, approximately 0.30% of the model. PEMA also substantially improves the LLaMA backbone and is competitive with rank-8 LoRA, which introduces low-rank updates within the language model. In our implementations, PEMA uses fewer trainable parameters and achieves about 1.71 to 1.90 training-time speedup over LoRA, although it does not reduce peak GPU memory. These results show that much of the task-specific performance improvement of a pretrained 3D medical MLLM can be recovered through a compact external adapter while leaving the underlying MLLM unchanged. Post-hoc controls show that similar Task 6 accuracy can be obtained without meaningful image information or with substantially simpler adaptation: zero-image training reaches 73.73% accuracy and a 3,072-parameter constant residual reaches 72.18%, compared with 73.74% for standard PEMA. These results show that high benchmark accuracy does not by itself establish strong use of the medical image.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.