acceptodds
Under review as a conference paper at ICLR 2027

RoboPhys: Beyond Outcomes, Toward Physically Correct Execution of LLM-Generated Robotic Programs

Abstract

Large language models (LLMs) are increasingly used to generate executable programs for robotic systems. As generated code moves from purely digital environments to embodied systems, however, conventional notions of correctness become insufficient. Existing code-generation benchmarks primarily measure functional correctness, considering a program correct if it executes successfully and passes functional tests. Yet robotic programs must also satisfy constraints imposed by physical execution. To capture this requirement, we formalize physical correctness as a distinct evaluation dimension that measures whether a program respects physical constraints throughout its execution in the environment. Importantly, outcome correctness does not necessarily imply physical correctness: a program may complete the intended task while violating physical constraints during execution. We introduce RoboPhys, an evaluation protocol that augments functional testing with physical execution constraints. RoboPhys identifies LLM-generated programs that satisfy terminal task predicates, then re-executes them under diverse physical conditions while monitoring constraint violations throughout their trajectories. Across five code-generation systems and 300 generated robot programs, 38 of the 136 nominal functional successes incurred at least one physical violation (27.9%). Trajectory audits guide task-conditioned repairs that restore task completion without executed violations or rejected hazardous requests in 42 of 57 initially unsafe-success program–scene cases. Cross-stage regressions motivate complete-execution revalidation and identify state-aware, task-agnostic repair as a direction for further study. These findings expose a blind spot in outcome-only evaluation, demonstrate the value of execution-level diagnosis for repair, and support evaluating physical execution correctness alongside task completion in generated robot programs.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.