MMCOMET: A Large-Scale Multimodal Commonsense Knowledge Graph for Visual Reasoning
Abstract
Commonsense knowledge is widely used to support multimodal reasoning and generation, yet existing resources are predominantly text-only and lack grounding in visual context. This limitation often leads to the retrieval of semantically related but visually misaligned knowledge, and can degrade downstream performance. We introduce MMCOMET, a large-scale multimodal commonsense knowledge graph that integrates physical, social, and eventive knowledge with visual grounding. MMCOMET extends ATOMIC2020 by associating each commonsense triple with representative images through a hybrid alignment pipeline, resulting in over 900K multimodal tuples spanning 19 relation types. We further establish a unified evaluation framework across diverse vision-language tasks, including visual commonsense reasoning, visual question answering, visual storytelling, image editing, and image captioning. Our analysis shows that multimodal grounding provides more contextually aligned and reliable knowledge compared to both text-only knowledge graphs and LLM-based knowledge generation, particularly in reasoning-intensive scenarios. At the same time, we observe that the effectiveness of external knowledge is task-dependent, and naive knowledge integration can introduce noise in perception-dominated settings. To mitigate this, we propose an adaptive knowledge router that dynamically selects optimal paths or bypass knowledge in harmful settings. MMCOMET provides a scalable multimodal commonsense resource and a benchmark for studying the role of grounded knowledge in multimodal systems.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.