Perceiving Before Cognizing: Agentic Visual Enhancement for Robust Document Parsing
Abstract
Modern document parsing models perform well on clean images but remain vulnerable to real-world visual degradations, motivating input enhancement to improve robustness across parsers. Yet document-dependent enhancement effects and potentially erroneous tool choices call for adaptive policies that can reassess intermediate results and explore alternative paths. We propose Perceiving Before Cognizing, which defines visual enhancement before parsing as an independently learnable interactive decision-making task, using parsing utility to guide optimization and improve the accessibility of task-relevant visual evidence. We introduce the AgenticDoc framework to support the reuse of this enhancement capability across parsers. To support policy learning for this task, we introduce Action-to-Token On-Policy Distillation (A2T-OPD), which bridges the gap between action-level teacher supervision and token-level generative policy learning. AgenticDoc improves all 13 evaluated parsers on Wild-OmniDocBench by 7.15 on average, including an average gain of 7.44 across nine parsers not used for training feedback. Overall performance on OmniDocBench v1.6 documents is preserved, with an average score change of +0.41. On PureDocBench, average gains across ten parsers reach 2.48 across the three scenarios.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.