acceptodds
Under review as a conference paper at ICLR 2027

IAR: Implicit Action Repetition for Policy Optimization with Sequence Critics

Abstract

In actor–critic learning, policy updates depend on both the learned value function and the action contexts in which it is queried. We study action repetition as a mechanism for constructing these contexts during policy optimization. We introduce implicit-action repetition (IAR), an actor-update mechanism for continuous control with sequence-conditioned critics. For each replay state, IAR samples an initial action and constructs prefixes that repeat it for different durations. The policy then samples a continuation action conditioned on each prefix, and the actor is optimized using the resulting sequence values. Repeated prefixes and critic parameters are held fixed during differentiation, allowing the actor to learn under hypothetical sustained-action contexts without requiring those repetitions to be executed in the environment. Once the initial action is sampled, all prefixes are available, allowing a causal Transformer to evaluate the continuation distributions in parallel. This removes sequential feedback among continuation predictions during actor updates, while environment interaction and critic bootstrapping retain autoregressive sampling. Across all 50 manipulation tasks under the Meta-World ML1 protocol, the resulting method achieves an aggregate success rate of approximately , compared with approximately for T-SAC. These results support repeated-action conditioning as an effective mechanism for policy optimization with sequence critics.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.