acceptodds
Under review as a conference paper at ICLR 2027

What the Label Admits, the Model Asserts: Evidence-Constrained Supervision of Vision-Language Models from Expert Prose

Abstract

Professional inspection practices hold years of photographs paired with expert free-text opinions written for human readers, not for model supervision. We study how to turn such legacy prose into supervision for a vision-language model (VLM) that produces structured, evidence-constrained defect reports, using a private corpus of 4,231 photographs from 586 expert-witness housing-condition reports prepared for disrepair claims in England. Each surveyor opinion follows an evidence–cause–remedy convention; a locally hosted text-only LLM converts only the visible-evidence component into a controlled JSON schema over a 52-type, 20-family defect taxonomy derived from the corpus, so that diagnosed causes never become visual labels, and every record passes structural, taxonomic and semantic validation, with flagged records excluded from training. Three open-weight VLMs of 7–9B parameters (Qwen2.5-VL, Qwen3-VL, GLM-4.1V) are adapted with QLoRA on a single 24 GB GPU under a property-level, near-duplicate-safe split. Adaptation raises type-level micro-F1 from at best 0.155 (few-shot prompting) to 0.432 with disjoint bootstrap intervals, and schema-valid output from 69.6% to 100%, while preserving the expert's evidence-only register. Under five-fold property-level cross-validation with a paired property-level bootstrap, swapping among three backbones from two vendors moves micro-F1 by at most 0.03, an order of magnitude less than adaptation itself, and the newest backbone lowers the no-defect false-positive rate from 56.0% to 44.0%. A label-content ablation then separates recognition from faithfulness: converting the whole opinion instead of only its evidence leaves recognition unchanged within bootstrap noise, but the model reproduces the register of its labels almost exactly, asserting an unverifiable cause on 95.6% of images against 0.0% for the evidence-only model. Recognition is set by adaptation and data volume (a training-set-size ablation shows no saturation); what the model asserts is set by what the label admits. We release the schema, taxonomy, prompts, validation rules and evaluation code; the imagery remains private.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.