acceptodds
Under review as a conference paper at ICLR 2027

SWAP: Switched Action Chunking Policies with Massively Parallelized Multi-Task Reinforcement Learning

Abstract

In massively parallel on-policy Reinforcement Learning (RL), performance saturates beyond a certain environment count because additional environments yield redundant trajectories due to executing similar actions sampled around the policy mean. Standard entropy regularization fails to address this, leading prior single-task methods to artificially partition environments across follower policies to force behavioral diversity with separate entropy bonuses. We propose that scaling *multi-task* reinforcement learning in this regime is unique in that policies inherently provide behavioral diversity by virtue of solving multiple tasks. The key insight is that agents can execute cross-task or *switched* actions during any task's rollout, using the otherwise excess rollout capacity of massively parallel simulation. However, utilizing this diversity to induce coherent exploration requires temporal commitment through action chunking. During a rollout, in a designated subset of parallel environments per task, policies replan by sampling action chunks across all task policies and executing the candidate that maximizes the value under the true task. By relying on pathwise policy gradients that differentiate directly through the critic, the policy absorbs this off-policy exploration without importance sampling correction. We call this framework Switched Action Chunking Policies (SWAP), which only incurs inference overhead during training that scales linearly with the number of tasks. In Meta-World, SWAP outperforms recent methods addressing performance saturation in the massively parallelized regime.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.