Seeing Is Not Judging: Evidence-Grounded Evaluation of Multimodal Models for Professional Preschool Environment Assessment
Abstract
Can multimodal models assess a real environment the way a trained professional does? Professional assessment asks a model to decide, from the complete visual record of a site, whether each written criterion of a rating scale holds, and to justify every decision with evidence. This requires finding the few relevant views among about a hundred photographs and applying the conditions of quantity, accessibility, and relation that each criterion states. Existing benchmarks either supply the relevant views or leave the criterion implicit, and so cannot tell whether a model failed to *see* the evidence or failed to *judge* it. We introduce **PEVA-Bench** (Preschool Environment Visual Assessment), a benchmark built on a professional rating scale: 6,450 expert decisions on 50 criteria over the complete photographic records of 129 classrooms, whose labels agree with independent onsite ratings on 91.0% of these decisions. Expert evidence and facts that can replace a model's own, together with criterion pairs that differ only in strictness, separate seeing from judging. Across fifteen models, the two fail in different ways. First, **Seeing separates models**: with expert facts, Qwen3.5-4B reaches 91.42 Macro-F1 alongside GPT-5.5's 90.46. Second, **Judging limits the strongest**: 80% of GPT-5.5's incorrect assessments involve criterion application, and 98% of these assign Satisfied to an Unsatisfied reference; no model tested tells a looser criterion from its stricter version in more than 42% of the cases that require it. PAJ (Perceive, Aggregate, Judge) structures visual evidence and criterion application, improving eleven of the thirteen models tested, most of all those limited by seeing, by up to 21 Macro-F1 points. *Seeing the evidence and applying the criterion emerge as distinct capabilities in professional assessment.* Code is available at [this anonymous repository](https://anonymous.4open.science/r/PEVA-Bench-68DE).
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.