acceptodds
Under review as a conference paper at ICLR 2027

Temporally Directed Gromov–Wasserstein Imitation Learning

Abstract

Cross-domain imitation learning studies how to transfer expert behavior between agents with different embodiments. Gromov–Wasserstein matching enables comparison across different state-action spaces without learning an explicit mapping, but its global coupling makes each imitation reward depend on the complete trajectory and ignores temporal ordering. We propose Temporally Directed Gromov–Wasserstein Imitation Learning (GWτdIL), which replaces full pairwise geometry with temporally ordered comparisons of local transition magnitude and directional variation. This produces per-transition imitation rewards from a single expert demonstration while preserving cross-domain structural comparison. For equal-length trajectories with uniform weights, the resulting empirical objective can be computed in linear time. We establish sufficient conditions for policy-optimality preservation and show that task and embodiment matching impose complementary constraints on the solution space. Experiments across domains with different state dimensions and morphologies show that GWτdIL consistently learns target-directed behavior and outperforms prior cross-domain imitation methods under comparable assumptions. Ablations isolate the roles of temporal ordering, directional information, and task–embodiment matching.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.