Can Audio-Capable Language Models Predict Human Emotion Judgments of Infant Vocalizations? Raw Audio Alone Is Not Enough
Abstract
Audio-capable large language models can process raw audio, but whether they reproduce fine-grained human emotion judgments for non-linguistic vocalizations is unclear. We compared four strategies on 178 matched transformed infant vocalizations rated for happiness, sadness, anger, fear, and disgust: direct GPT-Audio prediction from raw audio, supervised prediction from engineered acoustic features, and two feature-based in-context learning conditions using acoustically similar labeled examples with GPT-Audio and GPT-4o-mini. Direct GPT-Audio showed little correspondence with human ratings (r = -0.173 to 0.016), whereas supervised feature-based models reached r = 0.883. Feature-based in-context learning produced positive correspondence for both GPT-Audio (r = 0.450 to 0.725) and GPT-4o-mini (r = 0.573 to 0.786), although both remained below the strongest supervised models. These findings suggest that structured acoustic representations and labeled demonstrations support model prediction of human affective judgments, while conventional supervised feature-based approaches achieved higher correlations in this study.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.