SocialPrism: Multi-Dimensional Generative Reward Modeling for Video Social Reasoning
Abstract
Video social reasoning requires models to interpret dynamic, complex, and often incongruent multimodal signals to infer hidden mental states. Existing reward modeling approaches in reinforcement learning predominantly rely on blunt scalar scores, which collapse complex social reasoning processes into coarse feedback and make it difficult for models to identify or correct specific cognitive failures. To address this bottleneck, we introduce SocialPrism, a multi-dimensional and generative reward framework that decomposes holistic social judgments into five fine-grained cognitive dimensions. We construct SocialPrism-Data to provide multi-dimensional scores and rationales for reward model training, and SocialPrism-Bench, a human-verified benchmark for reward evaluation. Furthermore, we propose Gen-SocialPrism-7B, a generative reward model trained with a two-stage curriculum learning strategy. On SocialPrism-Bench, it achieves the lowest overall score prediction error among the compared reward models. It also outperforms standard binary-reward baselines in post-training and improves test-time reward-guided selection over standard sampling baselines.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.