Beyond Answer-Checking: Generalizable Verifiable Rewards from Task Understanding
Abstract
High-quality reward signals are critical for enhancing the capabilities of large language models on complex tasks through reinforcement learning post-training. However, existing reward approaches struggle to simultaneously achieve generality, verifiability, and interpretability. Learning-based reward models and rubric-based rewards often rely on additional models for semantic judgments, making the reward criteria and evaluation process difficult to explicitly verify, while verifiable rewards are mainly limited to tasks with easily automated evaluation and often provide only coarse-grained feedback based on final outcomes.To address these limitations, we propose Task-understanding-based General Verifiable Reward (T-GVR), a task-understanding-based general verifiable reward framework that explicitly transforms the model’s understanding of task requirements into verifiable supervision signals, providing fine-grained reward signals beyond final outcome evaluation.Specifically, T-GVR constructs a structured task representation containing task categories, task types, and key constraints, based on which it derives a Task Understanding Reward and employs a one-to-one constraint alignment strategy to improve the reliability of constraint reward estimation.Extensive experiments on mathematical reasoning, code generation, tool use, and open-ended generation tasks demonstrate that T-GVR consistently improves model performance across diverse task scenarios and highlights the effectiveness of task understanding rewards as general reward signals. Further analysis shows that explicit supervision over key task constraints is an important factor underlying the effectiveness of the Task Understanding Reward, while structured task representations and constraint alignment mechanisms further enhance the effectiveness and stability of the reward signals.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.