acceptodds
Under review as a conference paper at ICLR 2027

SoundSense: Evaluating Visual Sound Understanding in Comics and Webtoons

Abstract

Vision and sound provide complementary evidence for understanding physical events. While generative models increasingly produce audiovisual content and vision–language models (VLMs) are widely evaluated on image captioning and visual question answering, their ability to infer auditory events from static imagery remains less explored. We introduce SoundSense, an evaluation framework and annotation resource for studying whether depicted moments warrant sound effects. Comics provide a useful testbed because artists explicitly represent selected sounds through onomatopoeia. On eight Manga109-s volumes, we use these annotations as a proxy for sound placement, remove the lettering through local inpainting, and analyze VLM predictions across 6,238 panel-level decisions. Complementing this proxy, we provide 450 author-confirmed sound-effect regions across four vertical-scroll webtoons, including an audio-backed core of 359 moments with 421 reference clips. Among 444 regions with lettering metadata, 65.5% contain no sound-effect lettering. We benchmark five VLMs, including two hosted open-weight model families, on a frozen scope of 808 webtoon segments. We report localization, output-format coverage, cohort differences and clustered bootstrap intervals. On 750 shared parse-valid segments, localization F1 ranges from 0.159 to 0.680. An output audit reveals that coordinate-format compliance and spatial matching materially affect these scores, highlighting the need to distinguish spatial output failures from sound judgments. SoundSense provides a reproducible basis for evaluating visual sound understanding, analyzing prediction errors, and developing sound-aware systems for visual narratives. We release the webtoon annotations, reference audio, and source links.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.