acceptodds
Under review as a conference paper at ICLR 2027

MarineMammalBench: Sample Rate and Embedding Grid as Logged Variables in a Benchmark for Marine Mammal Self-Supervised Audio Encoders

Abstract

Self-supervised (SSL) audio encoders are now routinely reused far from the speech and general audio they were trained on. Two preprocessing choices then matter a great deal and are rarely reported: the input sample rate, and therefore the acoustic bandwidth the encoder retains, and the temporal grid on which it emits embeddings. We treat both as logged, swept experimental variables and study them in marine mammal passive acoustic monitoring, where the discriminative energy runs from killer-whale call types below 8 kHz to sperm-whale clicks above 20 kHz. We propose MarineMammalBench: a two-track protocol (in-domain SSL pretraining on a curated approximately 2.17M-hour unlabelled hydrophone pool; frozen-probe screening of public encoders under a matched protocol), a shared task ladder, and four mandatory controls, including a grid–tolerance coverage diagnostic. First-generation results give five findings. In a fixed 32 kHz sperm-whale click detector, restricting bandwidth to 3.5 kHz reduces event F1 from 0.7246 to 0.3999 at 20 ms while detector weights and the embedding grid remain unchanged. The historical 32–8 kHz F1 gap grows 2.8× when the matching collar tightens from 20 to 2.5 ms, so the temporal evaluation rule sets the size of the effect. For killer-whale call types the bandwidth gain is small (0.015 macro-F1 from the band above 8 kHz), so bandwidth has to be evaluated per task. On the official 31-class BEANS Watkins split, frozen marine animal2vec reaches 92.04% accuracy, ahead of Perch 2.0 (89.97%) and Whisper-tiny (87.91%); validation-selected readout raises Whisper accuracy from 75.22% at its final representation. On HICEAS minke-whale detection, AVES-bio instead leads the linear comparison with AP 0.5044, showing that encoder ranking depends on the task. On the prior system's presence test set, whether the model beats the human annotators depends on the metric and on the precision floor. The controls make bandwidth, temporal evaluation and uncertainty explicit parts of marine SSL benchmarking.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.