DepressionBench: A Multimodal Benchmark for Depression Recognition and Reasoning Analysis with Large Language Models
Abstract
Depression is a common and serious mental health condition, but its signs can be difficult to recognize in what people say, write, and do across everyday contexts. Large language models (LLMs) could help identify patterns in the language and behavior of people experiencing depression and explain the basis for their assessments. However, it remains unclear how reliably LLMs recognize signs of depression from text, audio, and video collected in different everyday contexts. We introduce DepressionBench, a curated large cohort datasets consists of 34,973 participants and a unified benchmark for depression recognition and analysis of model-reported reasoning. It covers text, audio, vision, and audiovisual inputs across four source scenarios and three languages. To assess depression-recognition performance and whether post-training improves models’ ability to identify signs of depression across modalities and everyday contexts, we evaluate twelve open-weight models under zero-shot and few-shot prompting, supervised fine-tuning, and reinforcement learning. Also, to examine whether the models’ explanations are consistent with the evidence in the inputs and to identify where their predictions fail, we assess the quality of their reported reasoning and analyze common prediction errors. DepressionBench provides a foundation for future research on multimodal depression recognition, including methods that generalize across everyday contexts. Our project page is available at: https://anonymous.4open.science/r/DepreesionBenchmark-C669
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.