acceptodds
Under review as a conference paper at ICLR 2027

When Does Sim-and-Real Co-Training Make Simulation a Reliable Evaluator?

Abstract

Evaluating robot policies in the real world is expensive. Simulation offers a scalable alternative, but its value depends on reliably predicting real-world performance. A common belief is that sim-and-real co-training can improve this predictiveness, but when and why remain unclear. We identify co-training's two opposing effects on sim-and-real agreement: sim-transferring carries real-learned behaviors into simulation, whereas sim-overwriting lets simulated data determine simulated behavior regardless of the policies' real-domain behavioral differences. We use controlled synthetic experiments to isolate how gap type, simulation-data dose, and policy-set construction govern these effects. We then test the predictions with manipulation policies across behavioral and visual sim-to-sim gaps and with real-robot evaluation. For behaviorally diverse policies, small simulation-data doses can improve behavior-wise agreement through sim-transferring. Larger doses can hurt agreement through sim-overwriting on states covered by simulated data, while continuing to improve agreement on Sim-OOD states, where direct sim-overwriting is avoided. Step-wise correlation between simulated and target-domain scores generally improves with sim-and-real co-training, but this does not necessarily imply behavioral alignment. These results provide practical guidance for reliably designing and interpreting simulation-based policy evaluation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.