MetroBench: Execution-Grounded Measurement for Visual Code Generation
Abstract
Visual code generation couples interpretation of a reference artifact with program synthesis, yet artifact fidelity and calibrated capability are distinct measurement problems. We introduce MetroBench, an execution-grounded framework and diagnostic study toward their joint evaluation. The framework specifies Construct, Reflect, and Repair protocols, and separates a four-component chart scorer from an independently audited psychometric layer. On 200 procedural chart items and 1,000 program variants, mean fidelity declines from 1.000 to 0.658 across five perturbation levels; 197 item profiles decrease strictly, while three exhibit local reversals. Removing plotted-data agreement reduces the strictly ordered profiles to 145, exposing sensitivity to the scoring rubric. In a separate synthetic-response experiment with 20 repetitions, 40 Fisher-selected items yield held-out ability RMSE 0.104 versus 0.110 for random selection and 0.090 for the 200-item full test; the paired 40-item gain is uncertain. A matched-bank study shows that anchor linkage can recover scale despite nearly unchanged rank correlation. However, nominal 95% plug-in ability intervals cover only 45.8% of held-out abilities, compared with 95.8% when item parameters are known. These findings demonstrate why execution grounding, adaptive efficiency, and uncertainty calibration require separate evidence. The present studies evaluate procedural scoring and synthetic measurement, not model capabilities; a validated common scale across multimodal perception and code synthesis remains an open empirical requirement.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.