acceptodds
Under review as a conference paper at ICLR 2027

RACE: Learning Multimodal Search through Referential Abstraction and Compositional Evidence

Abstract

Training multimodal search agents requires tasks that make visually grounded search necessary for determining the answer. Yet synthetic questions can admit correct answers while bypassing the image or ignoring evidence from intended targets, weakening the resulting supervision. We propose RACE, an image-conditioned task synthesis framework that constructs dependencies from images to targets and from target-specific evidence to answers. Referential Abstraction (RA) replaces identity-revealing text with specific visual references or ambiguous relational references resolved by searching among depicted candidates. Compositional Evidence (CE) links retrieved facts about multiple entities in one image or entities from different images into joint questions whose answers require every designated contribution. Both operations also apply to images found during search, allowing new visual observations to determine subsequent search targets or supply answer evidence. Fine-tuning Qwen3.5-27B on only 12K teacher-generated trajectories from RACE raises average accuracy across ten multimodal benchmarks from 53.5 to 56.0, a relative improvement of approximately 4.7% over the matched base agent. Our code, data, and model will be publicly available.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.