DocHarness-RSI: Failure-Guided Harness Evolution around Frozen Vision-Language Models for Long Multimodal Document Understanding
Abstract
Answering questions over long multimodal documents requires locating sparse evidence scattered across complex layouts and hundreds of pages, all within a fixed page and context budget. Frontier Vision-Language Models (VLMs) reason well once the right pages are in view, yet the harness that selects and organizes those pages is typically a static retriever or a heuristic truncation rule, leading to missed evidence, incomplete cross-page coverage, and diluted context. We argue that this harness, rather than the model, is the primary bottleneck, and that optimizing it around a frozen VLM unlocks substantial headroom with no parameter updates. We present DocHarness-RSI, which treats the context-construction harness as an executable program and improves it recursively. The inner harness implements a Retrieve-Complete-Purify procedure: multi-channel retrieval surfaces candidate pages, question-conditioned expansion completes cross-page evidence chains, and utility-based purification removes low-value pages so the budget is spent on what matters. The outer loop is failure-guided: it inspects residual errors on development data, traces each one to the responsible harness stage, mutates only that stage, and admits a candidate program only after it passes dual-gated verification. Because the search is anchored to the currently binding bottleneck, effort migrates across stages as earlier failures are resolved, which keeps the harness compact while steadily raising both evidence recall and context purity. On long-document understanding benchmarks, the evolved harness delivers consistent gains over strong retrieval and truncation baselines, showing that executable harness evolution is an effective optimization layer for frozen VLM agents. Code will be released.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.