acceptodds
Under review as a conference paper at ICLR 2027

SolarVQA: A Diagnostic Benchmark for Visual Reasoning in Photovoltaic Inspection

Abstract

Visual inspection of a solar cell involves more than detecting a defect. Questions about location, count and covered area require different properties to be combined. We introduce SOLARVQA, a diagnostic visual question answering benchmark with 106,555 questions over 16,322 electroluminescence images. Seven core tasks and six compositional tasks are generated from existing region annotations through image-integrity checks, consistent geometry, explicit ambiguity rules and controls for selected answer shortcuts. Across eleven zero-shot vision-language models, the highest task-macro score remains below a training-fitted nonvisual prior. However, supervised adaptation reveals substantial learnable visual signal. After one epoch of LoRA training, Qwen2.5-VL-7B reaches 64.49% test accuracy, compared with 41.64% before adaptation. With shuffled images, the corresponding scores are 38.02% and 37.08%. The benefit of correct image pairing therefore increases by 21.91 percentage points, with a paired 95% image-cluster interval of [20.86, 23.02]. Compositional performance also improves, although questions that change a simpler answer expose persistent failures. We further show that counting options can inflate ordinal agreement without improving exact accuracy. SOLARVQA supports both learning and diagnosis: its controls identify visual progress while revealing what aggregate gains leave unresolved.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.