Rethinking Feed-Forward Visual Prediction with Meta-Learned Adaptation Priors
Abstract
Feed-forward networks efficiently predict geometry, cameras, and appearance, but can lose accuracy under domain shift. Standard supervised training optimizes prediction without explicitly preparing models to benefit from limited target-domain supervision. We introduce MetaFeed, a framework that learns adaptation priors for supervised few-shot domain calibration. Through first-order meta-learning, MetaFeed optimizes a shared initialization for prediction after support-set updates to a task-dependent subset of existing weights. Geometry tasks adapt encoder attention projections and feed-forward layers, whereas appearance tasks adapt decoder convolutions. Calibration uses a small labeled support set once per target domain, after which the model processes held-out scenes with fixed weights and ordinary feed-forward inference. Controlled comparisons with supervised training and target fine-tuning under matched data and adaptation protocols show improved cross-domain prediction with competitive in-domain performance. Multi-source controls further show that the calibrated meta-trained model outperforms matched target fine-tuning on a domain unseen during source training. Experiments with VGGT reveal two distinct gains in reconstruction quality: meta-training reduces Chamfer distance before calibration relative to matched supervised training, and supervised calibration using a single support episode from a different scene within the same target domain further reduces it on disjoint target scenes.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.