The Gordian Knot for VLMs: Diagrammatic Knot Reasoning as a Hard Benchmark
Abstract
A vision-language model can fail a diagram question by misreading the diagram or by failing to operate on what it read. KnotBench pairs a corpus of 858,318 rendered diagrams of 1,951 prime-knot prototypes with 2,000 evaluation items in 14 tasks, labeled from the knot census and stored planar diagram (PD) codes. Its equivalence and move questions come as images and as PD codes, to separate these two sides of a perception-operation gap. We evaluate Claude Opus 4.7 and GPT-5 at two thinking settings each, under unequal output caps. Of the 56 (task, configuration) pairs, 22 (18 outside C1, whose baseline is 0) are at or below the uniform random baseline, and on 6 of 14 tasks no configuration scores above 65%. All configurations count crossings worse than a constant answer and match images to PD codes at chance. Pooled over the paired questions, they score 14.3 to 33.8 points higher on PD codes than on images. Thinking adds 2.0 points for Claude and 5.4 for GPT-5, mostly on three PD-code tasks, two of which text rules solve. The data establish a failure to read structure from our drawings; since real-step move labels also follow from crossing counts, whether the models simulate moves remains open.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.