acceptodds
Under review as a conference paper at ICLR 2027

What Does Binary Accuracy Prove? A Diagnostic Decomposition of Visual Graph Reasoning

Abstract

Vision-language models are increasingly used to read structured diagrams and reason over the relations they encode. Benchmarks typically score only the final answer. Ask a model "Is there a cycle?" and it almost always says yes. Ask "Which nodes form it?" and the cycle it names is often not in the graph. A correct answer thus cannot show whether the model recovered the diagram's structure. We use cycle detection as a diagnostic, since a named cycle can be checked against ground truth. We score four stages: which edges are extracted, whether a real cycle survives among them, whether the named cycle comes from the model's own list, and whether it is real. We test four models on cyclic graphs and trees, each drawn in a standard layout and a stress-test equal-angle ring layout. On ring drawings, models say "cycle" 84-100% of the time, yet the cycle they name is real only 4.8–44.0% of the time. Degree probes and human judgments suggest the drawings are not simply unreadable. Wrong answers follow two patterns: the real cycle never enters the extracted edge list (most common for three models), or it does and a fabricated cycle is named instead (one model). A clean text edge list restores two models fully but not the others, so extraction explains the gap only in part. A "yes" to "Is there a cycle?" does not show that the cycle a model names is real, so such errors can pass unnoticed and cause silent failures in graph-based workflows. The code and dataset are available at https://anonymous.4open.science/r/What-Does-Binary-Accuracy-Prove-31DB.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.