acceptodds
Under review as a conference paper at ICLR 2027

DR Pixel: Page-Native Deep Research with Vision-Language Models

Abstract

Complex documents such as scientific papers contain rich information in tables, figures, and equations; however, existing deep research (DR) agents typically reduce papers to OCR-extracted text. We instead ask whether vision-language models (VLMs) can enable page-native deep research: directly retrieving and reasoning over rendered document page images. We introduce DR Pixel, the first end-to-end multimodal DR agent that uses rendered page images as its sole retrieval and reasoning unit, operating over an open corpus of 1.06M pages. To evaluate this setting, we construct DRPixelBench, a benchmark on scientific papers requiring reasoning over text, figures, tables, and equations. Surprisingly, page-native DR already works well without training, outperforming OCR-based baselines across state-of-the-art models like GPT-5.4, Gemini-3.5-Flash, and Qwen3.6-27B, while requiring substantially fewer tool calls. We next turn to DR training. A key challenge is that demonstration traces are inherently noisy: some reasoning and action steps are critical to reaching the correct answer, while others are redundant or failed attempts, yet SFT treats them all equally. To address this, we introduce counterfactual credit reweighting, which weights individual reasoning and action steps according to their contribution to the final answer. Our training improves retrieval recall by about 20% and final accuracy by over 10% on DRPixelBench, and likewise outperforms the untrained baseline on all four additional benchmarks. On out-of-distribution ViDoRe-V3, DR Pixel comes within 1.8% of an oracle that is given the gold pages directly. Our results suggest that the conventional OCR-then-retrieve pipeline is not only unnecessary but suboptimal: directly retrieving and reasoning over visual document pages makes complex DR both more effective and more efficient.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.