What Language Makes Models See: Cross-Lingual Bias in Sensitive Judgments by Vision-Language Models
Abstract
Bias evaluations of vision–language models (VLMs) are usually conducted in a single language, leaving open whether sensitive judgments remain stable when only the prompt language changes. We evaluate five VLMs on the VisBias dataset using semantically aligned English, Spanish, and Chinese prompts across education, salary, politics, and religion. Under a controlled forced-choice setting, we distinguish cross-lingual consistency from within-language differences across ethnicity, gender, and occupation, including directional rank contrasts for the ordinal salary and education tasks. Cross-lingual behavior varies substantially by task and model: political orientation shows pronounced instability, although low valid-response coverage limits some comparisons, whereas salary is comparatively consistent. However, high agreement can partly reflect concentrated default responses rather than robust cross-lingual consistency. Demographic analysis reveals recurring group-conditioned patterns, including higher predicted salary ranks for male images in most model–language configurations and systematic ethnicity-related shifts. Overall, our results show that single-language evaluations can obscure both language-conditioned variation and demographic patterns in sensitive VLM judgments, motivating their joint assessment in multilingual bias evaluation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.