acceptodds
Under review as a conference paper at ICLR 2027

On the Structure of Reinforcement Learning Policies: Recovery from Synthetic Observations via Geometric Coverage

Abstract

Can a trained reinforcement learning policy be recovered without trajectories, rewards, or valid environment states? Across control tasks, students trained only on teacher-labeled synthetic inputs recover substantial closed-loop performance. Recovery is not explained by physical realism alone. Instead, it depends on two geometric properties of the query distribution: excitation of teacher-sensitive directions and sufficient variance outside the dominant sensitive subspace. Matched subspace interventions isolate directional excitation, while orthogonal thickening isolates ambient support; on Humanoid-v5, -scale orthogonal noise raises recovery from to with almost unchanged sensitivity overlap. We show theoretically that covariance mismatch governs transfer from query-space error to deployment error and that strict subspace queries can leave deployment behavior unidentifiable. A synthetic-only pilot further shows that inverse-covariance coverage can diagnose failures missed by raw overlap, although it is not universally predictive. These effects replicate for SAC and discrete-action ReLU policies. Small real-state buffers provide complementary deployment information, and action-imitation accuracy can continue improving after closed-loop return saturates. Together, these results identify geometric coverage as a central factor governing passive recovery of state-based policies.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.