AdAlign-Bench: A Benchmark for Human-Aligned Advertisement Subjective Video Understanding
Abstract
Subjective video judgments reflect viewers' experiences and perceptions, raising the question of whether multimodal large language models (MLLMs) align with human evaluations. Existing subjective benchmarks commonly use aggregated human ratings, which measure population-level alignment but offer limited insight into viewer differences and conditions of model–human divergence. We introduce AdAlign, a benchmark centered on population-level subjective rating alignment, using individual responses, psychological traits, and information-processing-related ratings to examine variation in model-human correspondence that aggregate scores may obscure. AdAlign comprises annotations by 225 participants on 393 advertisement videos, covering creativity and multimodal perceptual quality alongside open-ended impressions and self-reported viewing experiences. Current MLLMs align more closely with human ratings of creativity than of perceptual quality, capturing broad rating trends while exhibiting discrepancies in rating agreement and absolute scores. Alignment declines in both higher-rated and higher-disagreement video subsets, more markedly in the former. Under generic viewer prompts, model ratings also exhibit trait-related patterns of correspondence with human judgments. An exploratory case study suggests that correspondence with human information-processing reports may support creativity alignment, while illustrating differences between human and model response-level correlation structures. Together, these findings highlight variation in subjective rating alignment that aggregate performance alone may obscure.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.