acceptodds
Under review as a conference paper at ICLR 2027

BioInteract: A Large-Scale Multimodal Dataset for Evaluating Fine-Grained Semantic Understanding of Biotic Interactions

Abstract

While recent advances in vision-language models (VLMs) have spurred the development of domain-specific datasets and benchmarks, these evaluations often fail to assess fine-grained semantic understanding, allowing models to achieve high scores without robust visual grounding. We address this evaluation gap through the lens of biotic interactions: directional, asymmetric relationships between organisms (e.g., wasp parasitizes caterpillar vs. caterpillar parasitizes wasp). This relational complexity yields naturally adversarial instances that expose superficial reasoning in current VLMs. To this end, we introduce BioInteract, the largest multimodal dataset for evaluating VLM robustness on real-world biodiversity challenges. Curated from iNaturalist and validated against established interaction literature, the dataset contains 15.4K unique interactions spanning 6.5K taxa across 256K images. Each interaction is structured as a source-relation-target triplet, enabling controlled semantic perturbations. We further introduce BioInteract100, an image retrieval benchmark revealing that state-of-the-art VLMs suffer from severe consistency gaps and are highly brittle to relation-direction reversals. BioInteract provides a faithful evaluation of multimodal AI, while encouraging the development of robust systems for ecological research.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.