acceptodds
Under review as a conference paper at ICLR 2027

Response-Aware Exploration for Continuous Control Reinforcement Learning

Abstract

In reinforcement learning, exploration with sparse rewards and continuous action spaces on long-horizon tasks remains difficult, and a common strategy is to perturb the policy's actions during exploration. However, diverse action perturbations do not necessarily lead to diverse behavioral responses. Existing methods may encourage behavioral diversity indirectly, but few derive the perturbation directions themselves from the behavioral responses they actually produce. In this paper, we propose Discovered Action-Response Eigenbasis (DARE), a method that derives exploration directions by comparing perturbed and unperturbed rollouts from the same initial state and applies temporally correlated perturbations along those directions. DARE is reward-free, and is designed as an agile method that can be integrated into existing RL pipelines with minimal changes. Experimental results show that DARE outperforms representative action-space perturbation and temporally structured exploration baselines on sparse-reward tasks in Maniskill. DARE enables exploration to reach task-progress states more often, achieving up to 504% as many reaches as the strongest evaluated baseline, while remaining effective across different RL algorithms.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.