acceptodds
Under review as a conference paper at ICLR 2027

CAPSULE: Runtime Control and Hot Patching for Code-as-Policy Robotic Agents

Abstract

Code-as-Policy (CaP) robotic agents generate executable programs to orchestrate perception, planning, and control for robotic manipulation. However, relying solely on rollout-level feedback makes it difficult to associate execution failures with the relevant code segments. Recovery is further complicated by physical changes that persist after a failed execution, such as an upright book being knocked flat, which can make rerunning the program unsafe or ineffective. We introduce CAPSULE, a runtime control layer for CaP agents that supports group-level execution and effect-aware program repair without requiring the control layer itself to call existing robot APIs. CAPSULE partitions each program into source-local, effect-bounded execution groups and maintains a persistent namespace. It links execution traces and task feedback with the corresponding source regions to support inspection, local patching, and recovery. When recovery is needed after physical effects have occurred, CAPSULE avoids replaying groups whose physical effects have already occurred and generates forward recovery patches conditioned on updated environment observations. Across eight backbone LLMs on Robosuite and LIBERO, CAPSULE outperforms the CaP-X multi-turn baseline, boosting macro-averaged success rates from 36.7% to 49.7% and 8.2% to 16.8%, respectively. Ablation studies show that local patching and physical-state forward recovery provide complementary gains. Furthermore, we propose CAPSULE-Critique-GRPO (CC-GRPO), a post-training scheme that leverages execution-grounded repair trajectories to enhance standalone program generation. In a two-task study, CC-GRPO outperforms vanilla GRPO, increasing task success rates from 17.5% to 31.3% on Cube Stack and from 16.3% to 28.8% on Spill Wipe.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.