acceptodds
Under review as a conference paper at ICLR 2027

SeeTheFlow: Structure-Guided Flowchart Parsing

Abstract

Flowchart parsing requires vision-language models (VLMs) to recover not only textual and visual content but also the underlying topology. VLMs performance degrades sharply as diagrams become larger and more densely connected. Existing approaches primarily rely on visual reasoning or additional VLM inference, while the benefit of explicitly supplying externally extracted structure remains less explored. We introduce **SeeTheFlow**, a VLM-training-free framework that extracts an explicit structural prior from flowchart images using deterministic computer vision and OCR, then conditions a VLM on this prior through a lightweight, cost-aware router that intervenes only on diagrams likely to benefit. Across six VLM backbones, SeeTheFlow improves connection weighted F1 by up to 30 percentage points on large, structurally complex diagrams and by 11.2 points on real-world flowcharts, while invoking a second VLM inference on only roughly one-third of inputs rather than applying structural guidance unconditionally. We further introduce a 1,200-instance benchmark with controlled structural complexity, isolating diagram topology from lexical and rendering variation. Our results show that explicit, selectively applied structural evidence can improve robustness to structural complexity while avoiding the cost of additional inference on every diagram. Our code and data are publicly available at https://anonymous.4open.science/r/SeeTheFlow_ICLR-0502.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.