acceptodds
Under review as a conference paper at ICLR 2027

Conditional Policy Advantage: Changing the Expert or Test Physics Can Reverse Policy Rankings

Abstract

Does telling a robot controller the physics it faces, such as gravity, really help it handle changing physics? Across four simulated locomotion tasks, we show that the answer, and which architecture uses physics best, depends on two rarely reported choices: the expert whose actions the controller learns to copy, and the physical conditions used for testing. With training states fixed, changing only the expert reverses architecture rankings. With trained controllers fixed, changing only the test physics reverses whether physics inputs help. Even test sets drawn from one distribution rank architectures in opposite orders in 7 of 24 comparisons. These reversals are predictable from separate data, and inside a single controller the expert decides where accurate physics hurts. We call this dependence Conditional Policy Advantage and propose an Expert-by-Test Grid to report it.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.