acceptodds
Under review as a conference paper at ICLR 2027

MinerU-Diffusion v2: Bridging Autoregressive and Diffusion Decoding for Document OCR

Abstract

Diffusion decoding offers a new path to parallel generation for document OCR, but training from scratch leaves the capabilities of existing autoregressive parsers underused. We introduce MinerU-Diffusion v2, a 1.2B-parameter block diffusion document parser initialized from MinerU2.5-Pro. The design balances decoding parallelism and recognition quality while accounting for the downstream effects of layout decisions. Autoregressive-to-Diffusion Continued Training (ADCT) reuses existing document supervision through progressive block-size training and smaller-block replay, enabling a single checkpoint to support multiple inference block sizes. Hierarchical Joint Policy Optimization (HJPO) uses downstream rewards and conditional credit assignment to optimize final parsing quality, alternately updating a shared policy for layout generation and content recognition. On OmniDocBench v1.5, the model achieves Overall scores of 92.63 with predicted layouts and 94.82 with ground-truth layouts, outperforming standard MinerU2.5 in both settings. Compared with MinerU-Diffusion, it uses fewer than half the parameters while improving the respective scores by 3.69 and 1.45 points. Effective output throughput increases from 131.22 to 152.36 tokens per second, with flexible speed–accuracy trade-offs supported by configurable inference block sizes.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.