Zero-Training Feature-Space Alignment via Information Geometry
Abstract
Deep vision models often degrade under distribution shift. While test-time adaptation improves robustness by updating model parameters during inference, it typically requires iterative optimization, hyperparameter tuning, and multiple forward backward passes. We propose Zero-Training Fisher Geometry Alignment (ZFGA), a closed-form method that improves robustness under covariate shift without modifying model parameters. Our key insight is that distribution shift induces geometric distortions in feature space. ZFGA estimates the Fisher information matrix of the predictive distribution with respect to feature embeddings and applies a linear transformation that aligns the Fisher geometry of test features with a reference geometry computed from clean data. This can be viewed as a natural-gradient-inspired preconditioning step in feature space. We evaluate ZFGA on CIFAR-10-C and ImageNet-C using ResNet-50, DINO ViT-S/16, and CLIP ViT-B/32. ZFGA yields small but consistent improvements, with gains increasing as model robustness decreases. ZFGA is significantly better than zero-shot inference on all three models, but is not the strongest method on every individual model: covariance whitening yields a larger gain on ResNet-50, and Fisher whitening is statistically indistinguishable from ZFGA on CLIP. Comparing against six alternative training-free and gradient-based methods (covariance whitening, Fisher whitening, TENT, T3A, LAME, AdaNPC), ZFGA is the only one that is non-negative across all three model families, while every other method substantially harms at least one.A weak positive correlation between Fisher geometry distortion and ZFGA gain (Pearson , ) provides preliminary evidence that geometric misalignment contributes to robustness degradation. Requiring only forward passes and matrix operations at inference time, ZFGA offers a lightweight, deterministic, and reliably non-harmful alternative to optimization-based test-time adaptation, albeit with smaller gains on highly robust models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.