When the Simulator Fails, Not the Policy: Auditing Contact Artifacts in Simulated Robot Benchmarks
Abstract
Simulation benchmarks score robot policies by success rate, and every failed episode counts against the policy. Testing a vision language action model and a world action model in simulation, we find that failures in which the policy has placed the object and the simulator then ejects it are frequent, and that the benchmark's own scripted demonstrations fail the same way. We present SimuGuard, an auditing plugin for LIBERO, RoboTwin, RoboCasa and RoboDojo. It leaves the evaluation unchanged. At every physics step, it records the state and flags an object that starts moving at a contact faster than the robot's contact or gravity account for. It replays a recorded episode bit for bit, and it reruns the episode with one solver parameter changed. With it we trace each such episode step by step to the contact solver's response at a resting contact: within one physics step the solver reports an overlap of about two centimeters between the placed object and the thin wall of its container, removes the whole overlap in that same step, and the separation speed throws the object out at several meters per second. In SimuGuard's controlled reruns, a limit on the separation speed slows the ejection and lets most of the flagged failures pass. The artifacts change the score, and every way of scoring them rests on an assumption. Reported rates assume every ejection is the policy's fault. Replaying the recorded actions with the separation speed limited on every body from before each artifact assumes those actions remain the policy's choice once the ejection is gone. Excluding the flagged episodes assumes they are as hard as the rest. Running the policy closed loop with the limit on for the whole rollout assumes the limit changes nothing the task needs, and our data refute this one: it lowers the score of a task the artifacts barely touch. SimuGuard lets a benchmark separate what the policy did from what the solver did.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.