GeoTAP: Resolving Video Geolocation Ambiguity via Temporal Audio Prototypes
Abstract
Determining where a video was filmed is a fundamental challenge in spatial intelligence and computer vision. While modern visual foundation models excel at geolocating static images, adapting them to continuous video streams introduces two core failure modes: spatial aggregation noise, where naive per-frame averaging on the sphere degrades accuracy below a single static frame, and visual candidate ambiguity, where visually interchangeable scenes leave top geographic predictions tied with high confidence. To address both challenges, we propose GeoTAP, an encoder-agnostic framework that formulates video geolocation as hypothesis disambiguation within a unified Bayesian pipeline. GeoTAP extracts spatial candidate hypotheses via geodesic clustering and spherical geometric-median filtering to eliminate aggregation noise, then disambiguates candidate hypotheses against temporal audio patterns using learnable acoustic prototypes in our TAP (Temporal Audio Prototypes) module. Visual prior entropy naturally regulates audio influence during Bayesian posterior selection. Evaluated on our new AVG benchmark of 20538}diegetic-audio walking-tour clips across 907 sites, GeoTAP raises city accuracy from 27.9% to 34.2% and country accuracy to 69.4%, outperforming state-of-the-art visual and multimodal baselines without backbone retraining.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.