acceptodds
Under review as a conference paper at ICLR 2027

OutnumberedTales: Prevalence-Aware Evaluation of Contrastive Learning under Extreme Binary Imbalance

Abstract

Extreme class imbalance is a defining challenge across countless real-world domains where detecting rare events is critical. Despite this reality, the standard playbook for evaluating contrastive learning models relies heavily on artificially balanced test sets and symmetric accuracy metrics. Evaluating models in this artificial vacuum fails to reflect actual deployment conditions and obscures whether specialized, minority-aware algorithms provide any genuine benefit. Capturing a model's true efficacy demands a multifaceted approach. By examining performance across a broad spectrum of methods and decision-aware metrics, we demonstrate that different analytical lenses tell fundamentally different stories about model performance, exposing critical realities that conventional balanced testing actively conceals. Revisiting the problem through these varied lenses, we identify two distinct failure modes in existing minority-aware objectives. Geometrically, prototype-based contrastive losses degrade when fixed class prototypes conflict with the learned feature geometry. We introduce OrthoProto, an orthogonality hinge that resolves this conflict and surpasses the strongest published baseline wherever the failure is present, with gains scaling with the severity of the degradation. Statistically, batch-local objectives lose their supervisory signal when the expected number of minority samples per batch falls below one. We introduce CompMin, which maintains a census of minority representations and stays effectively invariant to batch size in this starved regime, a robustness we predict in advance from prevalence and batch size alone and then confirm. To ensure future research captures this multifaceted analysis, we release OutnumberedTales: a reproduction framework that implements published imbalance-aware objectives behind a common interface, pairs them with a decision-aware evaluation suite, and is designed for extension to new losses, datasets, and metrics.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.