TFoC: A Feature Gimbal for Radar Semantic Segmentation
Abstract
The proper utilization of spatio-temporal information is essential for Radar Semantic Segmentation (RSS). Classical 3D convolutions or spatio-temporal attention mechanisms can help an RSS model preserve spatio-temporal information to a certain extent. However, since the objects of interest in radar applications are dynamically moving, these "" computational modules are not the optimal choice in the absence of a spatio-temporal alignment mechanism. To address this, we propose TFoC (Temporal Focusing Convolution), motivated by the working principle of a camera gimbal, , dynamically adjusting the lens pose to steadily focus on the region of interest. Inspired by this, the receptive field design of TFoC aims to make spatio-temporal convolution operations focused rather than blind, and the sampling efficiency in the spatio-temporal domain dense rather than sparse. Therefore, TFoC is fundamentally a novel learnable module inspired by physical intuition. To better elucidate the learning mechanism of TFoC, we start from the classical Maximum A Posteriori (MAP) estimation and incorporate Bayesian theory to provide an approximate interpretive perspective for the forward computation logic and loss function design of TFoC.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.