Characterizing Age Conditioned Affective Reasoning in Vision Language Models
Abstract
Vision language models now support personalized and emotionally aware interactions, but demographic conditioning may reflect stereotypes rather than real human variation. We test this for age-conditioned affect using all 900 OASIS images and 17 proprietary and open-weight VLMs under no-age, young, middle, and older-adult conditions. We collect valence and arousal ratings, emotion labels, and open-ended descriptions, and compare them with participant-level judgments from 822 human raters. Models show high baseline competence (median valence ) and strong age responsiveness: 16 of 17 lower valence and all 17 lower arousal for older personas. Yet image-specific fidelity remains near zero ( to , ceiling ), even though model age effects agree with one another (mean pairwise ). The mismatch is content-dependent: on pleasant sensitive images, models lower older-persona valence by points while older human raters shift in the opposite direction (), and older-persona descriptions use more memory and vulnerability language. We also test six agent-based interventions on 300 images; reasoning alone does not improve alignment, while human grounding improves the broad direction of age effects but leaves substantial image-specific error, with the best agent reaching . These results show that matching average ratings or producing a plausible age persona is not enough: faithful personalization requires models to change on the same stimuli, and in the same direction, as the human group they aim to represent. Code and all model responses are available in the supplementary material.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.