Looking Where Infrared Matters: Instruction-Driven Instance Segmentation in Multispectral Vision
Abstract
Thermal imaging complements RGB under poor illumination, supporting perception in night-time driving, search and rescue, and industrial inspection. In these settings, instances of the same category can play different roles, and identifying the specific task-relevant ones requires descriptions of their behavior, function, or relation to the scene. Yet existing RGB–thermal segmentation produces semantic masks of category labels, leaving such distinctions unresolved. To bridge this gap, we introduce instruction-driven multispectral instance segmentation, where natural-language instructions specify which instances to segment. We therefore construct MINT, a benchmark of 5,000 registered RGB–thermal pairs with referring and reasoning instructions spanning single, multiple, all, and absent targets. Instructions express the desired target without mentioning thermal evidence, leaving the model to determine its relevance. This relevance varies across both instructions and spatial locations: helpful thermal cues coexist with harmful ones within the same scene. So, we develop a framework that learns when to consult thermal evidence and where to integrate it. Starting from RGB, a language model can emit a learned token to request regional thermal evidence and resume reasoning. An instruction-conditioned learnable gate then selects where to apply deeper thermal processing within the requested region. We supervise these decisions using intervention-derived utility estimates from a frozen full-evidence model, without manual utility annotations. Experiments on MINT show that our model outperforms both RGB-only and always-fused baselines, while processing thermal evidence in depth on only a fraction of each image.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.