acceptodds
Under review as a conference paper at ICLR 2027

Towards Recoverable Reward Transfer Geometry from IRL Gradients

Abstract

Inverse reinforcement learning (IRL) aims to recover reward functions from expert demonstrations, but reward ambiguity fundamentally limits their transferability across environments. Recent theory characterizes transfer sensitivity through principal-angle geometry between Potential Shaping Spaces, yet computing this latent geometry requires known or estimated transition laws. We show that this model-level route can be bypassed: we establish an observable route through IRL gradients, linking the transfer-relevant reward structure to gradient-space geometry and enabling direct estimation of this geometry. In practice, however, noisy or approximately evaluated gradients can lead to rank inflation and degeneracy of empirical principal angles, while complementary representations introduce structural angles unrelated to transfer. To enable estimation, we develop Dimension-Calibrated Truncated SVD (DC-TSVD), which uses intrinsic dimensions determined by IRL identifiability to suppress perturbation-induced directions and remove structural angles. Theoretical analysis establishes principal-angle estimation error bounds, and experiments on synthetic and finite-demonstration IRL environments demonstrate reliable estimation of the latent transfer geometry.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.