Tailor-Made for One: Learning from Each Test Pair for Zero-Shot Cross-Modal Homography Estimation
Abstract
Cross-modal homography methods typically rely on training data specific to the cross-modal combination encountered at deployment. Existing zero-shot methods alleviate this dependency by learning modality-invariant structural representations, but still treat each test pair as a passive input, leaving its pair-specific information unexplored. We introduce Tailor-Made for One (TMO), a zero-shot framework that turns each unlabeled test pair into an adaptation signal. The central challenge is that direct reconstruction-based adaptation can conflate appearance differences with geometric misalignment, allowing modality translation to distort the scene structure. TMO addresses this ambiguity by synthesizing prediction-centered training pairs with varying source geometries around the current alignment and requiring adaptation to remain effective across these pairs. This provides pair-specific supervision that refines appearance transfer while preserving its compatibility with a frozen geometric prior. On the three benchmarks shared with prior zero-shot work, TMO achieves state-of-the-art performance before adaptation. Across six benchmarks, adaptation tailored to each test pair further reduces MACE by 16.9%–50.8%. These results show that an unlabeled test pair is not merely an input to a transferable model, but a source of supervision for zero-shot geometric estimation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.