acceptodds
Under review as a conference paper at ICLR 2027

FloodMark: A Benchmark for Flood Scene Understanding and Reasoning in Multimodal Large Language Models

Abstract

Multimodal Large Language Models (MLLMs) are increasingly used to interpret complex visual scenes and support domain-specific decision making. To com- prehensively evaluate their capabilities in flood-scene imagery understanding, we introduce FloodMark, a benchmark for fine-grained, evidence-aware flood haz- ard assessment. FloodMark contains 1,400 real-world flood and post-flood images and over 16,000 human-annotated image question pairs across 15 expert-authored tasks, covering Evidence Quality, Scene Context, Flood-state Perception, Physical Reasoning, and Operational Reasoning. The benchmark evaluates whether mod- els recognize when visual evidence is insufficient to support a reliable conclusion. We evaluate 13 recent proprietary and open-weight MLLMs. Gemini 3.8 Flash achieves the highest overall accuracy of 73.2%, showing substantial room for im- provement. More importantly, when the image does not provide enough evidence for a reliable answer, on average, models still give a specific answer in 70.0% of cases. This tendency can lead to unsupported conclusions about flood conditions and their operational consequences. These findings show that current MLLMs remain unreliable for fine-grained, evidence-aware flood-scene assessment

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.