Visual Generative Foundation Models Are Reusable Priors for Non-Visual Sensors
Abstract
Visual foundation models have acquired transferable knowledge of scene geometry, spatial structure, and semantics from internet-scale imagery, yet sensors that remain informative where cameras fail, such as radio-frequency (RF) sensors, lack a comparable data ecosystem. Existing approaches either pretrain a separate foundation model for each sensor or distill one task at a time from a visual teacher, requiring substantial new effort for every new sensor or task. We ask whether a sensor can instead access the knowledge of a pretrained visual foundation model through a lightweight interface. In this paper, we present RFception, which maps RF measurements into the latent space of a pretrained visual generative perception model. Because the pretrained model selects tasks through text prompts, a single RF representation supports multiple geometric and semantic perception tasks without task-specific RF backbones or prediction heads. Across six held-out buildings, Reception outperforms task-specific RF models in depth estimation and, without task-specific RF supervision, performs surface-normal estimation, human pose estimation, and category-conditioned instance segmentation zero-shot. Its visual outputs can further be consumed by off-the-shelf vision-language models without fine-tuning, extending RF sensing to language-level reasoning. These results suggest that the knowledge captured by visual foundation models is not inherently tied to images, and that new sensing modalities can inherit this knowledge by learning an interface to the model rather than building a foundation model of their own.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.