DocAI: A Framework for Generating and Identifying AI-Edited Documents
Abstract
Document-level detection of AI-generated content is largely unaddressed, and off-the-shelf tools do not fill the gap: across nine AI-text and AI-image detectors, and both ways of composing them—score the whole page, or segment it and score each region—no configuration exceeds AUROC . The task demands dedicated tools and the data to build them; we provide both. DocAIgen edits real documents into large-scale corpora with region-level ground truth, from which we release a -page benchmark. DocAIdet is a document-native Vision-Language Model that reads the page as a whole, jointly detecting and segmenting AI-generated regions. It reaches page-level AUROC in-generation and – on two unseen generators ( at document level), with Dice and – respectively; the strongest piecewise baseline, retrained on our own data, reaches AUROC and Dice on the harder unseen generator, and the margins are significant with non-overlapping intervals. We release the engine, the benchmark and the detector.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.