WildBioWiki: Benchmarking Knowledge Use in Wildlife Question Answering with Controlled Inputs
Abstract
Wildlife question answering links animal identity to biological knowledge, but aggregate scores leave the benefits of supplied information and the sources of task difficulty unresolved. WildBioWiki aligns images with fixed species-level questions and evidence to support controlled comparisons. It contains 264,563 images of 392 vertebrate species, species-aware captions, and 3,136 unique question–answer pairs, inducing 2,116,504 image-conditioned QA instances. Across twelve zero-shot models, the best species-macro Acc@65 is 54.6%. Matched comparisons show benefits from correct names and further gains from retained knowledge. On balanced claims from 100 species, supplying identity and knowledge jointly raises decision-and-explanation accuracy by 44.0–55.5 percentage points across three models; human review on 20 species supports these gains. Matching required facts leaves an additional combination penalty unresolved, while human checks expose local errors hidden by high automatic agreement. WildBioWiki supports measurement of information benefits alongside checks on what benchmark scores establish.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.