acceptodds
Under review as a conference paper at ICLR 2027

AmbiHTMLBench: A Multidimensional Benchmark for Image-to-HTML Reconstruction of Heterogeneous Documents

Abstract

Document parsing systems often produce Markdown, which captures linear text but offers limited support in many scenarios (e.g., for fine-grained spatial structure, table geometry, visual hierarchy, and rendering semantics). HTML provides a richer representation for heterogeneous, layout-sensitive documents while supporting editing, inspection, and reuse. However, document-to-HTML reconstruction lacks a dedicated benchmark for systematic evaluation. Thus, we introduce AmbiHTMLBench to evaluate models' ability to reconstruct real-world documents as renderable HTML that preserves their content, structure, and visual appearance. AmbiHTMLBench comprises 1,115 diverse pages with structured annotations and two complementary evaluation protocols: (1) The objective protocol aligns rendered blocks between reference and predicted outputs to measure table fidelity, page-level text error, and two-dimensional layout consistency; and (2) the reference-anchored subjective protocol evaluates perceived fidelity using LLM judges calibrated against human preferences. Using AmbiHTMLBench to evaluate ten multimodal systems, we find different systems lead under the two protocols (with the highest objective score as 0.8092). We further demonstrate that task-specific fine-tuning enables a compact expert model to perform effective document-to-HTML reconstruction with 4.41 weighted points on Qwen-9B base.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.