acceptodds
Under review as a conference paper at ICLR 2027

Random Refresh Is All You Forgot: Rethinking Staleness in Offline GRPO

Abstract

Reinforcement learning for vision-language models is increasingly bottlenecked by rollout generation, motivating the reuse of stale, off-policy data. Existing methods address this problem through increasingly sophisticated selection, replay, reweighting, or clipping mechanisms, but rarely ask whether they outperform a much simpler alternative: refreshing a small random fraction of the rollout pool. We systematically compare eight methods that can be used to mitigate data staleness against a  15% random-refresh control across eight multimodal benchmarks. On the primary model family, none of 64 method-benchmark cells exceeds random refresh by more than 1.2 accuracy points, with a median margin of -4.2 points; 31 of 34 cross-family spot checks show the same pattern. We then ask where staleness actually concentrates. Across 26 family-benchmark cells and four model families, visually dependent tokens exhibit substantially higher staleness rates than language-prior tokens, with a mean high-|VD|/low-|VD| ratio of 9.1. Motivated by these findings, we construct a selective-refresh feasibility test that combines a small explorer with partial refresh by a larger trainee. The resulting system combines a small explorer with a larger trainee and refreshes groups with high visual dependence, supporting the first two findings while reducing the cost of full-size rollout sampling. Together, these results shift the focus from increasingly elaborate corrections of stale rollouts toward a simpler principle: maintaining freshness matters, and random refresh should be treated as a standard control whenever off-policy GRPO methods are evaluated.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.