GoLongRL: Capability-Oriented Long-Context Reinforcement Learning with Heterogeneous Multitask Alignment
Abstract
Reinforcement learning with verifiable rewards (RLVR) has shown promise for improving how large language models handle long contexts. However, existing long-context RL methods focus narrowly on retrieval-path-based data, resulting in limited task coverage and homogeneous reward formulations that fail to reflect the breadth of practical long-context requirements. We present a capability-oriented post-training framework that addresses this gap through two contributions. First, we construct an openly released dataset of 23K RLVR samples spanning 9 task types, each paired with its natural evaluation metric as the reward function. Under vanilla GRPO, training on this dataset alone outperforms comparable closed-source data recipes. Second, we propose TMN-Reweight, an advantage estimation method that combines task-level mean normalization for cross-task reward scale alignment with difficulty-adaptive weighting for more reliable gradient signals. Experiments on two model scales demonstrate that capability-oriented data is the primary driver of improvement, while TMN-Reweight yields additional gains on aggregation-intensive benchmarks. General reasoning, agentic memory, and dialogue memory capabilities are preserved or improved.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.