acceptodds
Under review as a conference paper at ICLR 2027

RLLR: Reinforcement Learning with Latent Reward

Abstract

Reinforcement learning with verifiable rewards (RLVR) has achieved strong performance in tasks with reliably verifiable answers, such as mathematical reasoning and code generation. However, extending RLVR to general reasoning tasks (\eg healthcare, engineering, law, and business) remains challenging, since correct answers are diverse in format and wording, making it difficult to design task-specific verifiers. Recent verifier-free methods instead derive rewards from the policy model's probability of generating a reference answer conditioned on the prompt and generated rationale. Nevertheless, we observe that probability-based methods are susceptible to non-semantic noise tokens and suffer from reference answer length effects, leading to unreliable reward evaluation. To reduce sensitivity to text format and wording, we shift from the text space to the latent semantic space to evaluate response quality. We propose Reinforcement Learning with Latent Reward (RLLR), a framework that extends RLVR to general reasoning tasks by defining rewards directly in the LLM's latent space without relying on external verifiers. Specifically, RLLR obtains the subspaces of the reference and generated answer's hidden states through Singular Value Decomposition (SVD) on answer hidden states, and measures their semantic consistency in subspaces. Experimental results demonstrate that RLLR improves the performance of Llama and Qwen models on five general-domain benchmarks and five mathematical benchmarks and outperforms rule-based, model-based, and probability-based reward methods.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.