RLAR: An Agentic Reward System for Multi-task Reinforcement Learning on Large Language Models
Abstract
Large language model alignment via reinforcement learning depends critically on reward function quality. However, generic reward models often underperform on heterogeneous task distributions due to distribution shifts, while training task-specific reward models is costly and prone to annotation difficulty, catastrophic forgetting, and loss of generalization. We present RLAR (Reinforcement Learning from Agent Rewards), an agent-driven framework that dynamically assigns tailored reward functions to individual queries. Specifically, RLAR transforms reward acquisition into a dynamic tool synthesis and invocation task. It leverages LLM agents to autonomously retrieve optimal reward models from the Internet and synthesize code-based verifiers through code generation. This allows the reward system to self-evolve with the shifting data distributions during training. On RewardBench-V2, RLAR outperforms static baselines and approaches the performance upper bound, validating the dynamic reward orchestration mechanism. In online RL post-training, RLAR yields consistent gains of 10%–60% across mathematics, coding, translation, and dialogue. Ablations confirm that synthesized code-based verifiers complement neural reward models on rule-heavy tasks, and that precise per-query reward routing outperforms monolithic reward prediction.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.