acceptodds
Under review as a conference paper at ICLR 2027

Leveraging Ecological Context in Biological Foundation Models at Scale

Abstract

Understanding the natural world often requires more than visual information: where and when an organism is observed provides critical evidence about its identity. Yet biological foundation models are trained predominantly from images and text, without exploiting readily available ecological context. We introduce a framework that incorporates geographic location and time of day and year directly into the semantic space of biological vision–language models. Using a novel dataset of more than 50 million multimodal observations, we train an ecological context encoder aligned to the BioCLIP 2 embedding space and an adaptive fusion mechanism combining ecological and visual evidence learned per-observation. Unlike other species-prior approaches, our method does not require a fixed species-specific classification head, preserving a flexible text-based retrieval interface in which candidate taxa are represented through their text embeddings. Across held-out species-recognition benchmarks, including novel biodiversity-weighted evaluation sets designed to better reflect biodiversity-rich regions underrepresented by observation effort, ecological context consistently improves recognition performance when combined with visual representations. Improvements are largest in geographically under-observed regions where visual training data are comparatively sparse. Our results demonstrate that large-scale ecological metadata provides a complementary modality for biological foundation models and offers a scalable route toward geographically robust biodiversity recognition.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.