Document Structure Beyond Heading Hierarchy: Dataset, Method, and Benchmark
Abstract
Documents organize meaning through their structure. Explicit structure helps retrieval and question answering, yet document parsing optimizes word-level accuracy and page-level fidelity while structure recovery targets the heading hierarchy alone. That formulation under-determines the structure of real documents: a parent need not be a title, an element may fall outside the ordered forest and rejoin it through a reference link to a target many pages away, and a unit need not be what one page yields, since page breaks split fragments and running headers take no part in it. We reformulate the problem as the joint prediction of a structure over a document: semantic elements that persist across page boundaries and carry a flag for taking part, a forest whose internal nodes need not be titles and whose groups may be named by virtual nodes, and directed reference edges from an anchor to a target anywhere in the document. We instantiate this with DocVine, a human-annotated dataset of 1,217 documents and 12,736 pages in which all three are annotated over the same elements. We design VinePipe, a staged pipeline, and train VineVLM-2B, a 2B vision language model fine-tuned to drive every stage of the pipeline. VinePipe with VineVLM-2B beats task-specific hierarchy parsers and frontier vision language models on DocVine, and leads on MPDocBench-Parse under its own protocol.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.