Free-to-Answer: Rethinking Output-Format Constraints in Reinforcement Fine-Tuning for Fine-Grained Visual Recognition
Abstract
Visual concepts in the real world are naturally organized from coarse semantic groups to increasingly fine-grained sub-categories, whose recognition requires both rich category knowledge and sensitivity to subtle visual differences. Rule-based reinforcement fine-tuning (RFT) has recently improved the fine-grained visual recognition (FGVR) capabilities of Multimodal Large Language Models (MLLMs). To enable automatic reward computation, existing methods commonly require responses to follow rigid output structures, such as separate reasoning and answer tags. However, it remains unclear whether these structures are merely parsing interfaces or whether they also affect recognition behavior. In this work, we systematically investigate how output-format constraints affect FGVR before and during reinforcement fine-tuning. Across five benchmarks, removing mandatory structural tags generally improves recognition accuracy, although the effect varies across models and datasets. Motivated by this observation, we develop Free-to-Answer, which optimizes free-form responses using LLM-based label extraction followed by deterministic verification. Under Open-Ended QA evaluation, Free-to-Answer achieves higher average accuracy than format-constrained RFT across the 1-, 2-, and 4-shot settings, while controlled experiments suggest that the improvement is not solely attributable to the reward design. We further analyze how free-form responses interact with closed-set label grounding. Our results reveal that rigid output structures can affect recognition behavior in visual RFT, rather than serving only as interfaces for reward parsing.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.