SPINE: Token-Selective Test-Time Reinforcement Learning with Entropy-Band Regularization
Abstract
Large language models (LLMs) and multimodal LLMs (MLLMs) excel at chain-of-thought reasoning but face distribution shift at test-time and a lack of verifiable supervision. Recent test-time reinforcement learning (TTRL) methods derive label-free pseudo-rewards from self-consistency voting over sampled trajectories, yet they often collapse: the majority-vote reward prevails, responses shorten, and Pass@1 declines. We trace this to uniform sequence updates in which most tokens are low-entropy followers, while a small high-entropy subset determines the reasoning branches. Thus we propose SPINE, a token-selective test-time reinforcement learning framework that (i) performs distribution-aware forking-token selection to update only decision-critical branch points, and (ii) applies a robust entropy-band regularizer at those tokens to prevent premature collapse and suppress noisy drift. SPINE plugs into GRPO-style objectives (optionally with a KL anchor) and requires neither labels nor reward models. Across eight benchmarks spanning multimodal VQA and text-only reasoning, SPINE consistently improves Pass@1 over TTRL while mitigating response-length collapse and stabilizing training dynamics. Experiments on 8B-scale LLM and MLLM backbones further demonstrate that these gains persist with increased model capacity, supporting token-selective updates at chain-of-thought branch points as an effective label-free mechanism for test-time adaptation. Code will be released.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.