Cog-Research: Mitigating Grounding Drift in Long-Horizon Multimodal Deep Research
Abstract
Multimodal deep research agents can acquire reliable evidence yet fail to preserve it through long-horizon reasoning and report synthesis. We characterize this failure as Grounding Drift: previously established visual observations, claim–source bindings, or task constraints become omitted, corrupted, or misbound at later stages. We introduce Cog-Research, a post-training framework that converts intermediate research evidence into supervision for grounded report synthesis through source-linked, task-specific rubrics constructed from research traces, workspace artifacts, and human reports. We build a process-grounded dataset of 500 Plan-and-Act research tasks and perform trajectory SFT followed by rubric-guided GRPO with rewards for evidence-informed criterion coverage and report quality. Built on Qwen3-VL-8B-Instruct, Cog-Research improves the base agent by 4.85/3.54/4.91 points on MMDeepResearch-Bench, DeepResearch-Bench II, and MiroEval-Multimodal, respectively, and reduces the task-level incidence of primary Grounding Drift errors from 30.7% to 17.1% on MMDeepResearch-Bench.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.