acceptodds
Under review as a conference paper at ICLR 2027

TCCA: Target-Conditioned Credit Assignment for Test-Time Reinforcement Learning

Abstract

Test-time reinforcement learning derives supervision from self-generated answers, but outcome-only credit loses information at two levels: within a prompt, responses with the same binary reward receive identical base advantages, while across prompts, within-prompt normalization does not reflect differences in target ambiguity. We introduce Target-Conditioned Credit Assignment (TCCA), which keeps a single hard pseudo-target estimated from the rollout pool and uses it to organize learning credit at two complementary resolutions. Self-certainty-weighted Soft Vote estimates the target and its answer distribution. The associated ambiguity reallocates normalized outcome credit across prompts through Hardness-Oriented Advantage Scaling, while target-matching reasoning traces form compact semantic references that provide relative reasoning credit within each prompt. The two terms are normalized separately and combined at the advantage stage for the same policy update, reusing the sampled responses and a frozen semantic encoder without a dedicated process reward model. Across single-run evaluations on three backbones and four reasoning benchmarks, TCCA yields higher point estimates than outcome-only TTRL in all 12 settings. In matched three-seed MATH-500 experiments, TCCA improves sample accuracy / majority correctness from 38.96/40.40 to 41.10/42.80 on Qwen2.5-0.5B and from 39.80/41.73 to 42.82/44.67 on Llama3.2-1B-Instruct, with positive paired differences in all three seeds on both backbones. Component removals and a separate one-versus-two-prototype comparison further support the credit construction, with approximately 5.3% additional complete-run time on the measured setup.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.