CoralAgent: From Visual Recognition to Constrained Agentic Ecological Assessment for Coral Reef Conservation
Abstract
Reliable coral reef ecological analysis requires more than accurate visual recognition; it demands strict adherence to sampling protocols, taxonomic standards, and precise measurement definitions. However, existing vision-language models and general-purpose agents lack explicit mechanisms to enforce these domain-specific requirements, a challenge compounded by the scarcity of reliable coral-specific models. To bridge this gap, we introduce CoralAgent, a multi-agent framework that translates ecological requirements into explicit rules for evidence collection, computation, and validation. Its core component, Coral Structured TAsk Representation (C-STAR), formalizes measurement objectives, evidence requirements, validation criteria, recovery policies, and admissible conclusions. Guided by C-STAR, CoralAgent coordinates reusable analytical primitives, validates intermediate results, and switches to workflows adapted from conventional coral surveys when coral-specific models are unavailable or their outputs fail validation. To evaluate quantitative ecological analysis beyond recognition and segmentation, we introduce Coral-Bench, comprising 1,986 question–answer pairs and 2,535 images across 23 tasks. These tasks cover coral recognition, visual knowledge, cover estimation, community composition and structure analysis, and ecological index computation, including cross-site and cross-year analysis. Experiments with seven multimodal baselines show that CoralAgent outperforms the GPT-5.4 baseline on 22 of 23 tasks and achieves the best or tied-best performance on 20. Compared with GPT-5.4, CoralAgent improves task accuracy for single-image coral cover estimation from 20.00% to 72.59% and for Gini–Simpson diversity calculation from 19.18% to 61.64%.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.