acceptodds
Under review as a conference paper at ICLR 2027

Theory-Driven Conceptualization and Implementation of Evaluation Metrics for Chart Editing

Abstract

Charts can be visually similar yet communicate different information, while visually different charts may support the same interpretation. This makes chart comparison fundamentally different from conventional image similarity and poses considerable evaluation challenges in tasks such as chart editing and chart generation. We define this as a matter of functional equivalence: whether a candidate chart preserves the information and inferences supported by a reference chart. To support the study of functional equivalence, we introduce ChartFun, a benchmark of human judgments on the functional equivalence of 2,400 model-generated charts against rule-based targets. ChartFun spans five chart types and six image- and code-based chart-generation models. Our analysis reveals substantial variation across models: only 60.5% of generated charts are judged functionally equivalent to their rule-based targets overall, despite many outputs being visually similar. More importantly, correlation studies reveal that existing evaluation metrics used in chart-generation studies show limited agreement with human judgments of functional equivalence. We explore three new evaluation approaches grounded in prior theory on data visualization: a vision-language model (VLM) judge, a structured theory-driven comparison pipeline, and a lightweight learned evaluator. Our best-performing method, the VLM Judge, reaches a Pearson correlation of 0.60. Our work opens new avenues for evaluating chart generation and editing through the theory-grounded framework of functional equivalence.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.