Beyond Aggregate Accuracy: A Deployment-Aware Dataset and Benchmark for Mixed-Defect Wafer-Map VQA
Abstract
Industrial diagnosis with vision-language models is hierarchical: decide whether a part is abnormal, name every co-occurring defect, localise each one, and explain it. We release Mixed-WM38 VQA, a benchmark that links these stages: 190,075 question-answer instances over 38,015 wafer maps and 38 normal, single-defect and mixed-defect combinations, with answers derived deterministically from labels or pixels, fixed splits, and a full-corpus audit that removed a shortcut in which location was a relabelling of defect type. The benchmark is built to expose what aggregate scores hide. Only 2.66% of the evaluation wafers are normal, so an always-abnormal gate is 97.34% accurate, and on a 4B backbone every standard recipe we tried (SFT, minority replication, domain-QA pretraining, chain-of-thought, three GRPO rewards) sits at that score while rejecting all 212 normal wafers and then narrating a coherent, image-unsupported diagnosis. The failure is not perceptual: a compact CNN separates normal from abnormal wafers perfectly on the same set, whereas in the VLM the correct "No" is the second-ranked token, about 18 times less likely than "Yes". Threshold calibration does not repair it: the false-alarm-minimising threshold still misses 52% of defective wafers on held-out data. The training interventions that do move the gate, a margin loss and a selective GRPO reward, trade false alarms for fatal misses in a seed-dependent way: one recipe yields 7.6% or 30.6% fatal misses depending only on the seed, and another declares every wafer normal in two of three seeds. A prompt that looks curative under teacher forcing (74% correct-answer probability) produces 94% fatal misses under generation. Cross-stage consistency is no safeguard either: chain-of-thought raises internal agreement to 87-100% while jointly correct gate-type-location tuples fall to 8.6%. We also audit the benchmark against itself: a classifier-to-template pipeline beats the VLM on every metric, which bounds how much of the protocol needs open-ended reasoning and gives future work a mandatory floor. Aggregate accuracy, class-conditional risk, cross-stage consistency and end-to-end correctness are distinct quantities that can move in opposite directions; hierarchical VQA systems should be selected on all four, with denominators and run-to-run variation reported.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.