FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning
Abstract
Reinforcement learning (RL) improves the reasoning capabilities of large-scale models. Group Relative Policy Optimization (GRPO) and its variants alternate between rollout sampling and policy update, updating the policy using sampled rollouts and their advantages. Unlike supervised learning with explicit ground-truth targets, these policy updates rely on high-quality rollouts as an implicit "teacher" to guide their direction. However, GRPO and similar algorithms sample only from the original prompt. For tasks beyond the policy's current capability, high-quality rollouts are rare, leaving updates without a meaningful gradient direction and causing training to stall. To address this issue, we propose FBOS-RL, a Feedback-Driven Bi-Objective Synergistic Reinforcement Learning framework. Specifically, we let the model perform Feedback-Guided Exploration Enhancement based on the feedback provided by the environment, and on top of this we design two mutually reinforcing training objectives: Exploitation-oriented Policy Alignment (EPA) and Exploration-oriented Capability Cultivation (ECC). Extensive experiments show improved rollout efficiency and higher peak validation performance. We match the total number of generated training rollouts with GRPO: GRPO samples 72 responses per original prompt, whereas FBOS-RL allocates the same budget to 8 initial and 64 feedback-guided responses. On TravelPlanner with Qwen3-8B, FBOS-RL achieves a relative gain of in peak final pass rate within the shared rollout budget. To reach GRPO's best final pass rate, GRPO requires approximately as many training rollouts as FBOS-RL. Additional experiments compare against On-Policy Self-Distillation (OPSD). FBOS-RL also exhibits higher policy entropy and lower gradient norms than GRPO throughout training. Code will be made publicly available.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.