Evaluating Vision-Language Models on Images of Individuals with Limb Deficiencies
Abstract
Strong performance on general benchmarks does not establish whether vision- language models (VLMs) correctly interpret images of individuals with limb deficiencies. For this population, applications such as preparing visual evidence for Paralympic athlete classification depend on correctly identifying residual limbs, prosthetic devices, and body-device configurations. We find that VLMs can miss this evidence or assign it to the wrong body regions, so later reasoning begins from incorrect visual information. To examine these failures, we introduce INCLUSIVEVLM-LIMBEVIDENCE, a test-only benchmark on about 17k open- world images of this population. It provides over 52k direct evidence queries and over 20k compatibility instances, with human reference results and prompt paraphrases. The design separates recognizing evidence from using it in struc- tured decisions. LIMBEVIDENCE-PRESENCE tests evidence detection. Because detection alone does not establish where evidence belongs, LIMBEVIDENCE- ATTRIBUTION tests body-region assignment in parallel Constrained and FreeForm formats. To examine whether models can also use this evidence in structured deci- sions, LIMBEVIDENCE-COMPATIBILITY tests option matching under benchmark- defined relations through parallel Category and Diversity probes. Across widely used general-purpose VLMs, models often detect limb evidence but fail to iden- tify its body region or side. Partial option matches often fall short of complete compatible-set recovery. Models can also repeat the same incorrect answer across paraphrases. The larger model in our scale comparison improves coarse detection but does not resolve attribution and set-recovery errors. These results reveal a gap between recognizing limb evidence and using its spatial and relational details correctly. They also show that response consistency cannot replace correctness. INCLUSIVEVLM-LIMBEVIDENCE makes these overlooked capability limits mea- surable and supports research toward visual AI that better serves individuals with limb deficiencies.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.