RIMAR: Riemannian Metric Attention and Regularization for Training-Free Open-Vocabulary Semantic Segmentation
Abstract
Training-free open-vocabulary semantic segmentation converts frozen vision-language features into pixel predictions. Existing inference methods typically use different geometric quantities for feature comparison and spatial refinement. We introduce RIMAR, which uses a Riemannian metric derived from spatial variation in a multi-layer feature field as a common geometry for both. Bidirectional Affine Attention (BAFAT) fits closed-form affine query-key mappings using normalized metric-area weights. Its area-weighted proposal is fused with native attention before value aggregation. The retained metric supplies directional distances to a refinement graph alongside RGB differences. Semantic information from the input and overlapping-window integration produce observations refined by fixed-budget Gaussian and heavy-tailed updates derived from maximum a posteriori (MAP) objectives. RIMAR requires no segmentation training, auxiliary visual backbone, or camera calibration. For a continuous BAFAT-MAP reference family, we establish reparameterization equivariance of scores and MAP solution sets under consistent transport of features, sampling measures, and regularization fields. With frozen CLIP ViT-B/16, RIMAR averages mIoU and pixel accuracy across 13 panoramic, fisheye, and perspective benchmark settings, exceeding PEARL by and percentage points, respectively, under a common evaluation protocol. Paired projection and sampling studies show higher mean accuracy and prediction agreement than PEARL and NACLIP, alongside residual sensitivity. Controlled comparisons show an upstream advantage over PEARL under a common refiner and a -point mIoU improvement over a jointly simplified geometry-independent variant.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.