SHARING ENVIRONMENT RANDOMNESS: WHEN REWARD VARIANCE FALLS BUT GRADIENT RISK RISES
Abstract
Group-based reinforcement learning is widely used to train language model agents by comparing rewards across sampled trajectories. Sharing environmental randomness can make these comparisons less noisy, but it also reduces independent environment coverage under a fixed rollout budget. We construct counterexamples with identical reward distributions but opposite sharing effects on gradient risk, showing that reward statistics alone cannot determine when sharing is beneficial. In this paper, we introduce Environment-Sharing Risk Analysis (ESRA), a framework for evaluating environment sharing through its effect on gradient estimation risk. We derive an exact finite-group risk criterion for raw leave-one-out gradients at a frozen policy, quantifying the trade-off between within-environment fluctuations and differences in mean gradients across environments. We test the criterion with independent gradient audits and controlled sharing interventions. Using Qwen3.5-35B-A3B-Base, we evaluate ESRA on HotpotQA and ALFWorld, covering multi-hop question answering and interactive decision-making. At matched total training cost, environment sharing alone improves native GRPO by 7.10 percentage points on average. Holding the raw-LOO estimation and calibration rules fixed, shared collection improves the two-task mean by 7.00 percentage points over independent collection.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.