acceptodds
Under review as a conference paper at ICLR 2027

Vision-Language Models on Specification-Governed Instrument Approach Charts: Which Failures, and Which Supplied Information Changes Them

Abstract

Converting engineering drawings into structured data requires recovering every specification-defined instance and its attributes, and recent vision-language models are a candidate solution. On instrument approach charts, they omit instances or fill in wrong attributes. Scoring whole records as right or wrong cannot tell an unformed instance from a nonconforming field set or a misread value, nor which information to supply. From charting specifications and chart evidence, we define the instances and attributes each chart requires; this shared semantic reference scores outputs and yields the required instances (instance list), the fields to fill (field list), and reading rules. We evaluate Luna on 862 charts and Luna, Qwen, and K3 on a common 171-chart subset. Most strict failures are required instances that never form a record; with the instance list, matching is almost fully restored, nearly all newly correct records are previously unmatched targets recovered whole, and matched instances show almost no net change. Unformed instances account for 85.28% and 92.16% of Luna's and Qwen's strict failures; with the list supplied and identities still self-written, matching rises from 45.7%, 12.2%, and 56.7% to 98.8%, 73.9%, and 98.9%. The field list mainly aligns output fields with the scored reference fields; its effect on unsupplied values depends on response design. The reference makes each failure type separately measurable, turning an accuracy gain into an account of which information changed what.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.