Open-Text Underwater Detection: Grounding Rich Descriptions Across Objects and Context
Abstract
Underwater object detection becomes challenging when object appearance alone is insufficient to distinguish target instances, while relations and environmental context provide complementary evidence. We study open-text underwater detection, where rich descriptions of appearance, relations, and context are used to localize all matching instances. We introduce AquaGround, to our knowledge the largest benchmark for open-text underwater detection, containing 224K images, 893K object instances, 2,363 categories, and 7.52M query records. AquaGround provides four progressively enriched query levels and five structured environmental attributes, with each query explicitly mapped to all valid matching instances. We further propose AquaDetect, a cue-guided open-text underwater detector that combines target semantics, underwater context, and detail–redundancy cues to guide progressive visual aggregation while preserving spatial correspondence for localization. Experiments on AquaGround show that AquaDetect reaches 58.94 AP, outperforming the SOTA method by 5.22 points.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.