TARS: Extending Reinforcement Learning Gains to the Long Tail under Limited Data for Domain-Specific Visual Recognition
Abstract
Visual recognition in domain-specific applications, such as biodiversity research and agriculture, often suffers from limited and imbalanced training data. Large Vision-Language Models (LVLMs) offer a promising starting point with rich visual and semantic knowledge acquired through web-scale pretraining, but naive supervised fine-tuning (SFT) on such data can cause them to overfit to head classes. Although reinforcement learning (RL) is reported to generalize better than SFT, our analysis shows that its gains under Group Relative Policy Optimization (GRPO) do not consistently extend to tail classes. To address this limitation, we propose **Tailness-aware Advantage ReScaling (TARS)**, a novel RL-based method that amplifies GRPO advantages for classes that are under-represented and poorly recognized. This amplification is guided by *tailness*, defined as the product of class rarity in the training data and the per-class error rate of the SFT model. Experiments on four domain-specific classification and detection benchmarks show that TARS consistently improves tail-class performance over SFT and vanilla GRPO, while outperforming existing imbalance-aware GRPO variants in both tail-class and overall performance. TARS also provides further gains when applied on top of imbalance-aware SFT methods, demonstrating that the two approaches are complementary.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.