acceptodds
Under review as a conference paper at ICLR 2027

How Fast Should a Model Commit to Supervision? Training Reasoning Models on the Tsallis Loss Continuum

Abstract

The standard recipe for training reasoning models — SFT on rationale annotations followed by RLVR — relies on rationale annotations being available and clean. When no annotated rationales exist, SFT cannot start at all, and when the model is too unaligned, RLVR alone stalls at cold start; when annotations are weak, SFT memorizes the errors. Using the Tsallis -logarithm, we define a loss family that interpolates between RLVR (at , the _exploitation pole_) and the log-marginal-likelihood over latent trajectories (at , the _density-estimation pole_). The standard pipeline corresponds to a stepwise schedule, and intermediate enables training in regimes where the two-step pipeline cannot apply. All members share the same gradient direction, differing only by a per-instance amplification that reweights each instance independently of the learning rate. Under gradient flow analysis, we show that the exploitation pole requires time to escape cold start but is robust to label noise, while the density-estimation pole escapes in but memorizes label noise. This separation explains why SFT () precedes RLVR () in the standard pipeline. We further derive two Monte Carlo estimators that directly optimize fixed- on the continuum, without annotated rationales or external verifiers: Gradient-Amplified RL (GARL) and Posterior-Attenuated Fine-Tuning (PAFT), with shared bias but different variance and stability properties. On FinQA, HotPotQA, and MuSiQue, GARL at sufficiently high escapes cold start where GRPO fails entirely under Qwen 3 and Gemma 4 architectures. Without task prompts, cold-start GARL on Qwen 3 0.6B also matches or exceeds prompted GRPO with reshaped reward on majority vote and best-of- on all three benchmarks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.