More Evidence to Build Than to Answer: The Sufficiency-Necessity Gap in Task Synthesis for Multimodal Search Agent Post-Training
Abstract
Multimodal deep search agents are increasingly post-trained on complex synthetic tasks designed to demand iterative retrieval and multi-step reasoning across visual and textual web evidence. Existing synthesis methods typically organize interconnected evidence into multi-step solution structures, deriving questions via structural constraints and solver-based filtering to promote task difficulty. However, evidence sufficient to construct an answerable query is not necessarily required to solve it: visually grounded facts may be recovered through text alone, intended premises can be bypassed, and intermediate answers may leak into the prompt. A task can therefore remain valid yet fail to supervise its intended capabilities—a mismatch we term the sufficiency-necessity gap. To bridge this gap, we introduce EvidenceWeave, a task synthesis framework that jointly addresses evidence access, premise dependence, and information exposure. By composing qualified atomic relations into multi-step dependencies while keeping intermediate answers implicit, EvidenceWeave ties evidence acquisition and intermediate reasoning to the information requested by the query. Extensive experiments show that our queries yield lower direct-answer recovery rates than the public queries evaluated and elicit visual evidence acquisition and use in correct teacher solutions. Fine-tuning Qwen3.5-27B on the resulting trajectories improves performance across five core benchmarks and leads to more varied and connected evidence use. These results support EvidenceWeave as an approach to narrowing the sufficiency-necessity gap by turning construction evidence into targeted supervision for the acquisition, interpretation, and use of multimodal evidence. Our code, data, and model will be publicly available.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.