acceptodds
Under review as a conference paper at ICLR 2027

Entropy Targeting in Sequential Reinforcement Fine-Tuning of Language Models

Abstract

Sequential reinforcement fine-tuning asks a language model to keep improving as its tasks change, yet the auxiliary signals added to stabilize such training are rarely checked for whether they reach the policy update at all. We study this question in AdaCoT-CL, a group-relative sequential learner that combines three consistency signals (entropy targeting on the reasoning span, alignment between reasoning and answer representations, and answer agreement) with replay, and we reanalyze 63 recorded runs on a five-task reasoning stream, 59 of them with Qwen3-8B. We first show that the agreement signal supplies no direct policy gradient: its group-shared reward bonus cancels whenever rewards are centered within each prompt group, and its auxiliary loss is detached. This result partitions the seven ablation variants into four effective objectives and motivates a grouped comparison. In that comparison, configurations with entropy targeting improve on average and those without it decline. On fixed monitoring panels, the 12 runs with entropy targeting gain 13.2 accuracy points from initialization to the end of the sequence, while the 9 runs without it lose 8.7 points (difference 21.8 points, exact permutation ; a contrast averaged within seed and alignment blocks gives 24.1 points, ), and the primary comparison replicates this direction against five baselines. The largest divergence occurs between the end of the LogiQA phase and step 32 of the final task, a block that combines training on the new task with replay of earlier prompts: mean accuracy on the earlier panels rises by 27 points on average with entropy targeting, and the smallest such rise, 16.7 points, exceeds the largest without it, 4.2 points. The evidence has clear boundaries: the panels are drawn from training prompts, initial accuracy is low under a 512-token completion budget, the benefit shrinks under other task orders, and the same objective collapses Llama-3.1-8B-Instruct. We therefore present entropy targeting as the consistency signal that distinguishes improving from declining configurations of this learner, and we specify the held-out protocol needed to establish it as a continual-learning method.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.