Reinforcement Learning toward Target Behavior
Abstract
Desired model behavior defines a destination for post-training. Reaching this destination through fixed rewards requires anticipating how output scores will shape model behavior. We instead make the joint satisfaction of behavioral requirements the explicit training objective, introducing Reinforcement Learning toward Target Behavior (RLTB). Specifically, RLTB formalizes these requirements as a joint target region in behavior space and minimizes the discrepancy between the policy's current behavior and this region. Each behavior's deviation from its target determines the direction and strength of its corrective feedback: the feedback weakens as the target is approached, can reverse after overshoot, and vanishes upon satisfaction. This feedback scores sampled outputs for training, yielding target-dependent output credit. We implement RLTB through Target-Behavior Group-Relative Policy Optimization (TB-GRPO), which converts this credit into group-relative advantages for GRPO's policy update. To prevent outputs from influencing the feedback used to score them, we estimate each prompt group's feedback from other independent groups using leave-one-prompt-group-out cross-fitting. In mathematical problem solving and factual question answering, models trained with TB-GRPO jointly approach heterogeneous behavioral targets, achieving the smallest joint target gaps among the evaluated baselines. Across the evaluated requests and base policies, the trained models approach their behavioral destinations without case-specific hyperparameter retuning. Together, these results position RLTB as a principled foundation for target-directed post-training.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.