acceptodds
Under review as a conference paper at ICLR 2027

AdjuvantChart: Benchmarking Multi-Panel Relational Understanding in Real-World Scientific Charts

Abstract

Scientific chart understanding requires models not only to read numerical values but also to identify experimental conditions and integrate evidence across panels to draw conclusions. Existing benchmarks focus primarily on visual information extraction, numerical reasoning, and single-turn scientific question answering. They place less emphasis on how models organize and use chart evidence in real experimental contexts. We introduce AdjuvantChart, a scientific chart understanding benchmark built from published adjuvant research papers. Guided by domain experts’ needs in literature analysis, we design three tasks: structured data extraction, experimental relation verification, and cross-panel composition. For the latter two tasks, we construct question families grounded in traceable atomic facts, comprising local, compositional, and condition-variation questions. We evaluate both answers and supporting evidence to assess whether models consistently use the same underlying chart evidence across related questions. AdjuvantChart contains 600 high-quality composite figures from published papers. The experimental relation verification and cross-panel synthesis tasks include 380 question families and 1,134 question-answer pairs. In our evaluation, GPT-5.5 achieves the highest exact-match accuracy of 81.71% on the question-answering tasks, while Gemini-3.7-Flash achieves the highest RMS-F1 score of 54.92% on structured extraction. Joint evaluation of answers and supporting evidence changes the relative ranking of some models, indicating that evaluating final answers alone does not fully capture models’ ability to use evidence from scientific charts. AdjuvantChart provides a new testbed for evaluating multimodal models’ chart understanding and evidence use in experimental contexts.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.