acceptodds
Under review as a conference paper at ICLR 2027

SocialPrism: Multi-Dimensional Generative Reward Modeling for Video Social Reasoning

Abstract

Video social reasoning requires models to interpret dynamic, complex, and often incongruent multimodal signals to infer hidden mental states. Existing reward modeling approaches in reinforcement learning predominantly rely on blunt scalar scores, which collapse complex social reasoning processes into coarse feedback and make it difficult for models to identify or correct specific cognitive failures. To address this bottleneck, we introduce SocialPrism, a multi-dimensional and generative reward framework that decomposes holistic social judgments into five fine-grained cognitive dimensions. We construct SocialPrism-Data to provide multi-dimensional scores and rationales for reward model training, and SocialPrism-Bench, a human-verified benchmark for reward evaluation. Furthermore, we propose Gen-SocialPrism-7B, a generative reward model trained with a two-stage curriculum learning strategy. On SocialPrism-Bench, it achieves the lowest overall score prediction error among the compared reward models. It also outperforms standard binary-reward baselines in post-training and improves test-time reward-guided selection over standard sampling baselines.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.