acceptodds
Under review as a conference paper at ICLR 2027

DocAgentic-R: Learning Utility-Aware Retrieval for Document Deep Research

Abstract

Training retrievers with downstream utility has improved agentic search, but existing frameworks assume homogeneous text passages. We study the same idea when candidates are text-carried document units that differ in modality and granularity. method trains a BGE-M3 retriever with local relevance (lr), a prefix-freeze suffix-replan answer-contribution score (gac), optional chain and localization labels, and multi-positive contrastive learning. On document-disjoint BlueprintSynth-Eval (2,430 questions; single runs), the full stack raises full-chain@20 from 0.588 to 0.674 and MRR from 0.713 to 0.800, above frozen Qwen3-Embedding-4B at 0.605. lr-only training falls below the backbone; gac is the signal that restores it. On the same pool, simpler positive rules reach 0.678–0.681, and text-only positives fall to 0.657. External transfer is about one point: MMDocIR Page R@5 moves from 0.705 to 0.710 and Layout R@5 stays below the backbone; the best LongDocURL setting gains 1.7 R@1 points and 1.0 LLM-judged point. Figure-unit coverage does not rise. Utility supervision helps synthetic multi-hop chain recovery. It does not, by itself, beat simpler utility labels or a native visual encoder.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.