acceptodds
Under review as a conference paper at ICLR 2027

AdAlign-Bench: A Benchmark for Human-Aligned Advertisement Subjective Video Understanding

Abstract

Subjective video judgments reflect viewers' experiences and perceptions, raising the question of whether multimodal large language models (MLLMs) align with human evaluations. Existing subjective benchmarks commonly use aggregated human ratings, which measure population-level alignment but offer limited insight into viewer differences and conditions of model–human divergence. We introduce AdAlign, a benchmark centered on population-level subjective rating alignment, using individual responses, psychological traits, and information-processing-related ratings to examine variation in model-human correspondence that aggregate scores may obscure. AdAlign comprises annotations by 225 participants on 393 advertisement videos, covering creativity and multimodal perceptual quality alongside open-ended impressions and self-reported viewing experiences. Current MLLMs align more closely with human ratings of creativity than of perceptual quality, capturing broad rating trends while exhibiting discrepancies in rating agreement and absolute scores. Alignment declines in both higher-rated and higher-disagreement video subsets, more markedly in the former. Under generic viewer prompts, model ratings also exhibit trait-related patterns of correspondence with human judgments. An exploratory case study suggests that correspondence with human information-processing reports may support creativity alignment, while illustrating differences between human and model response-level correlation structures. Together, these findings highlight variation in subjective rating alignment that aggregate performance alone may obscure.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.