CARE: CONSERVATIVE POLICY OPTIMIZATION FOR RLVR IN GENERATIVE CLASSIFICATION
Abstract
Reinforcement learning with verifiable rewards (RLVR) has become the predominant approach for post-training large language models on tasks where output correctness can be algorithmically verified, and GRPO-family optimizers (GRPO, DAPO, DR.GRPO), all of which inherit PPO’s clipped surrogate, dominate this space. Their design assumptions of high-entropy action spaces and well-dispersed rewards are inherited from multi-step mathematical reasoning and open-ended generation. Generative classification, where the policy emits a class label directly as a short generated string, satisfies neither: the action distribution is low-entropy, and while every rollout is scored, the reward takes only a handful of distinct values. We identify two temporally distinct breakdowns of this clipped surrogate under these conditions. Early in training, elementwise surrogate minimization selects the unclipped branch for negative-advantage high-ratio (NAHR) samples, admitting exactly the aggressive updates that clipping exists to suppress. Later, as accuracy rises, prompt groups become uniform and mean-centered advantages vanish, starving the policy of gradient signal. We propose Conservative Advantage Ratio Enhancement (CARE), which addresses both with two modifications and no new hyperparameters. First, the elementwise minimum between clipped and unclipped surrogates is replaced by a single batch-aggregated magnitude comparison, committing the whole update to the less aggressive surrogate and closing the per-sample escape route. Second, advantage normalization is applied conditionally, falling back to the raw reward when the batch carries no reward variance. We further show that the batch-level conservative surrogate selection enables entropy recovery in collapsed-policy regimes without an explicit entropy bonus. Trained on RAGTruth QA for hallucination classification, CARE surpasses GRPO by absolute +11.5 points macro-F1 on in distribution test set over three seeds, and transfers to OOD hallucination classification on RAGTruth (Summarization, Data2Text), HaluEval (QA, Dialogue, Summarization), and Internal QA (Easy, Hard). CARE attains 64.40 macro-F1 on the multi-class ANLI benchmark against 59.20 for the strongest baseline.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.