Rethinking Long-Tail Learning in Vision-Language Models: From Single Distribution to Modality-Coupled Distributions
Abstract
Vision-language models (VLMs) have achieved remarkable progress in multimodal processing and generation, yet still exhibit systematic failures in object hallucination and rare-concept understanding. Prior studies have shown that the concept distribution in multimodal data is markedly long-tailed, resulting in imbalanced downstream performance. However, existing long-tail methods mainly employ a single frequency statistic to define long-tail distribution, which cannot accurately reflect the effective long-tail distribution in multimodal scenarios. For example, the same concept may exhibit markedly different scarcity across modalities, leaving some concepts are frequent in text but weakly grounded visually. Moreover, shared multimodal representation spaces entangle long-tail biases across modalities, making it difficult to decouple modality-specific effects, causing incomplete rare-concept expression. We refer to this problem as modality-inconsistent long-tail. To address this, we propose a pre-decoupled multimodal-aware long-tail learning framework (PMLL), which introduces modality-decoupled reweighting to model multimodal long-tail distributions, rare-word reinforcement to improve rare token expression, and adaptive task-aware optimization to balance multi-task contribution. Furthermore, a keyword-decoupled long-tail evaluation metric is introduced to discover the detailed learning performance of informative tokens, while a metric-general co-occurrence-aware long-tail mining protocol is developed to reduce co-occurrence-induced bias in estimating long-tail distributions. Overall, we provide a unified framework for understanding, mitigating, and evaluating modality-inconsistent long-tail effects in VLM learning. Extensive experiments demonstrate the effectiveness of the PMLL under modality-inconsistent long-tail.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.