acceptodds
Under review as a conference paper at ICLR 2027

SEM-VQA: Structured Visual Question Answering for Scientific Understanding of Scanning Electron Microscope Images

Abstract

Understanding scanning electron microscope (SEM) images is central to materials science, yet progress on applying multimodal large language models (MLLMs) to this domain is limited by the scarcity of large-scale, open-ended visual question answering (VQA) datasets grounded in real scientific imagery. We introduce SEM-VQA, a pipeline for constructing and evaluating SEM-focused VQA data that enforces image-grounded reasoning. Using LLM-assisted annotation and prompting constraints requiring answers to be derivable from the image alone, we build a corpus of 224,026 question-answer-visual-evidence triplets over 47,981 SEM micrographs, organised into four levels of visual understanding: observation, detection, identification, and interpretation. We construct a held-out benchmark of 300 images via embedding-based clustering and use it to evaluate open-source and proprietary MLLMs under an LLM-as-judge protocol. Fine-tuning a compact 2B-parameter Qwen3.5 model on our corpus raises its judged score from 3.33 to 3.83 (5-point scale), a 14.9% relative gain that closes roughly 76% of the gap to GPT-5.4-mini (3.98). The fine-tuned model significantly outperforms Claude-Haiku-4.5 and, by a narrower margin, Gemini-3.1-Flash, though a modest but significant deficit remains relative to GPT-5.4-mini; it also leads all evaluated systems across every automatic metric (token-F1, BERTScore-F1, STS, ROUGE-L, METEOR, BLEU). These results show that targeted, structured dataset construction, rather than model scale alone, can substantially close the gap between small open models and large proprietary systems on specialised scientific visual reasoning. We release our dataset, benchmark, and evaluation protocol.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.