acceptodds
Under review as a conference paper at ICLR 2027

ChunkRL: Learning Adaptive Action Horizons for Fast and Accurate Robot Manipulation

Abstract

Robotic manipulation is not uniformly difficult: free-space motion can tolerate long open-loop execution, while brief contact-rich stages demand frequent re-observation and correction. Yet modern Vision-Language-Action models commit to a fixed action-chunk horizon, forcing every stage to share the same trade-off between redundant policy evaluations and delayed feedback. We present ChunkRL, which turns the action horizon from a static inference hyperparameter into an observation-conditioned feedback policy. ChunkRL first builds a parameter-efficient bank of short-, medium-, and long-horizon experts from offline demonstrations while freezing the backbone. It then optimizes an Adaptive Chunk Selector with per-decision Group Relative Policy Optimization, allowing the selector to infer feedback needs from closed-loop rollouts and trade off task success against feedback cost. We introduce Adaptive Layer-Pruning, which couples temporal feedback allocation with per-call computation by conditioning policy depth on the selected horizon and VLM residual signals. Across LIBERO, VLABench, and CALVIN, our 0.5B model improves average success rates by 0.8, 6.1, and 3.2 percentage points, respectively, while delivering – inference speedups. These results show that scheduling when to observe and how much to compute improves manipulation reliability and efficiency.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.