acceptodds
Under review as a conference paper at ICLR 2027

RoboBranch: Branching Exploration as a New Scaling Axis for Robot-Use Agents

Abstract

General-purpose AI agents are moving from browsers and terminals to robot bodies. These robot-use agents perceive through cameras, act through navigation, perception, and manipulation tools, and improve without any weight updates by learning task guidance from interaction. However, existing agents learn from independent rollouts, which reveal only one outcome at each state they visit: testing another action means replaying a similar trajectory and reaching a state that differs in both context and action. This raises a question: can a robot-use agent learn from many outcomes of the same state, rather than one outcome of many states? We answer it with RoboBranch, the first framework that makes branching exploration in simulation a new scaling axis for robot-use agents, featuring three designs: (i) Shared-State Branching, which executes multiple candidate actions from identical copies of each decision state; (ii) Outcome-Grounded Guidance Learning, which revises task guidance from every outcome, selected or not, successful or failed; and (iii) Width-Scalable Runtime, which proposes all candidates in one response and executes branches in parallel, so that wider exploration yields more outcomes without a proportional increase in wall-clock time. Across LIBERO-Pro, RoboCasa365, and RoboTwin C2R, RoboBranch improves over Harness VLA built on the same open-source Qwen model by 10.6 points on average and brings the open-source agent into the range of proprietary agents built on Claude Code and Codex. We show that simulation can do more than run additional trajectories: it lets an agent learn more from every state it reaches.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.