Text Assisted Event Stream Visual Place Recognition via Multi-Scale Spatial-Temporal Optimal Transport
Abstract
Benefiting from the high dynamic range, high temporal resolution and low energy consumption of event cameras, low‑altitude economy research leveraging event cameras has become a popular topic. Several existing attempts introduce natural‑language descriptions into visual place recognition (VPR) to mitigate the sparsity of structural spatial information inherent to event cameras. Nevertheless, these methods still fail to fully address two core challenges: effective encoding of event streams and semantic alignment between event streams and text. In this paper, we propose a multi‑scale spatial‑temporal optimal transport approach termed MSTOT‑EPR to tackle the above challenges. Specifically, given event streams and scene text information, we first adopt a visual encoder and a text encoder to extract respective feature representations. Optimal‑transport constraints are then applied to achieve cross‑modal alignment. Afterwards, a cross‑modal bottleneck fusion module is designed for further bimodal feature fusion. Furthermore, we develop multi‑scale spatial optimal transport and inter‑frame temporal optimal transport modules to mine high‑quality spatial‑temporal event features for scene recognition. Extensive experiments on the NYC‑Event‑VPR and EPRBench datasets fully validate the effectiveness of our proposed framework for event‑based VPR. The source code will be released upon acceptance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.