CDMMT: Conditional Diffusion with Multi-Modal Guidance for Multi-Object Tracking
Abstract
Because data modalities differ substantially, existing multimodal tracking methods often design specialized mechanisms for specific modalities, making it difficult to use multimodal information effectively within a unified framework. To address this issue, we propose Conditional Diffusion for Multimodal Multi-Object Tracking (CDMMT). Through modality-adapted condition encoding and injection, CDMMT incorporates different forms of auxiliary information into a shared paired-box denoising process for object localization and adjacent-frame association. Modality information is first encoded as a condition and injected into paired-box prediction. We further introduce energy-guided sampling: differentiable conditional energies measure the agreement between predicted boxes and modality information, and their gradients correct intermediate box predictions so that the constraints propagate to later predictions through sampling-state updates. Using Global Navigation Satellite System (GNSS) positions and infrared images, we demonstrate the encoding, injection, and guidance of sparse positional and dense image information. Experiments show that CDMMT effectively uses multiple modalities and improves tracking performance. Code is available anonymously at https://anonymous.4open.science/r/CDT-B163.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.