acceptodds
Under review as a conference paper at ICLR 2027

VisAudit: Evaluating Multimodal Agents for Visual Diagnosis and Repair

Abstract

Multimodal agents are increasingly applied to data visualization tasks, but remain limited in autonomous review: unlike human reviewers, they may fail to recognize when a visualization is incorrect, determine what should be changed, repair it without disrupting correct content, and verify whether the intervention succeeded. Existing benchmarks largely evaluate predefined individual capabilities such as chart generation, instruction-guided editing, or defect detection, and therefore do not capture this gap in autonomous review. We introduce VisAudit, a benchmark for evaluating the diagnosis, repair, and verification of autonomous visualization. Given a rendered chart and configurable auxiliary evidence, including its source data table, intended text summary, and visualization code, an agent iteratively diagnoses potential defects, modifies and executes visualization code, inspects execution and visual feedback, and determines when no further intervention is needed. VisAudit defines three tracks spanning diagnosed repair, autonomous repair, and open-world verification, and contains 1,900 flawed instances across 21 chart types and 10 flaw categories, together with 300 initially correct charts. We construct the benchmark through controlled perturbations of validated source visualizations, with systematic verification and human-aligned quality control to ensure that injected defects are well-defined and recoverable from the available evidence. Experiments with leading multimodal models reveal a substantial gap from reliable autonomous review: the strongest evaluated model fully recovers only 47.4% of flawed charts in the autonomous-repair setting. By exposing failures across diagnosis, intervention, preservation, and verification, VisAudit characterizes the capabilities needed to close the gap between current agents and human-like visualization review and provides actionable insights for designing more reliable autonomous multimodal systems.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.