CTRL-Z: Code agent Trial robot actions and Reason with Latent world model
Abstract
Code-as-policy approaches translate language instructions into executable programs, enabling compositional task planning and reuse of robot skills. However, generating a valid program don't establish whether a candidate action will produce the desired physical change. Action-conditioned world models offer a way to anticipate such changes, but their latent predictions are not directly expressed in the language of executable programs generated by code agent. To address these questions, we propose a closed-loop system in which a code agent can read, reason, make prediction and update object physical parameters based on those predictions, and generate actions accordingly. To make it readable for the agent, we build Structured Object Program, a compositional program describing object primitive geometry, physical parameters, interaction properties and rendering object programs into geometric views. For the prediction system, we propose a dual-stream world model that accepts RGB images and rendered geometric views as input while predicting future visual and geometric latent variables. Physical parameters regarding the dynamic parts of the object are decoded from these predicted latents and fed back to the agent, enabling it to understand how its current actions affect the object's future properties. Experiments across manipulation tasks and object generalization settings show that our framework improves task success, cross-object generalization, and agent interaction efficiency
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.