CheXpercept: A Benchmark for Evaluating Expert-Level Lesion Perception in Chest X-rays
Abstract
The evaluation of vision-language models (VLMs) for chest X-ray (CXR) analysis has largely been limited to disease-presence classification without visual grounding. Such evaluations fail to verify the expert-level lesion perception necessary to ensure the clinical reliability of VLMs. To address these limitations, we introduce CheXpercept, a sequential, multi-level perception benchmark that mirrors a radiologist's cognitive workflow across coarse-level detection, fine-level contour evaluation and revision, and semantic-level attribute extraction. To ensure high clinical fidelity at scale, we construct the dataset using a semi-automated generation pipeline paired with review by six medical experts. CheXpercept contains 12,000 QA items derived from 2,400 CXRs, covering eight clinically critical pulmonary and cardiac lesions. To demonstrate the current landscape of VLM perception, we benchmark 18 general and medical VLMs on CheXpercept. The models achieve adequate performance only at the coarse level, with accuracy degrading precipitously on deeper visual tasks. Notably, recent general models outperform recent medical VLMs overall. While medical models show marginal improvements over their specific backbones at certain perception levels, the inherent capability of the backbone itself ultimately dictates performance far more than current domain adaptation strategies.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.