Temporally Directed Gromov–Wasserstein Imitation Learning
Abstract
Cross-domain imitation learning studies how to transfer expert behavior between agents with different embodiments. Gromov–Wasserstein matching enables comparison across different state-action spaces without learning an explicit mapping, but its global coupling makes each imitation reward depend on the complete trajectory and ignores temporal ordering. We propose Temporally Directed Gromov–Wasserstein Imitation Learning (GWτdIL), which replaces full pairwise geometry with temporally ordered comparisons of local transition magnitude and directional variation. This produces per-transition imitation rewards from a single expert demonstration while preserving cross-domain structural comparison. For equal-length trajectories with uniform weights, the resulting empirical objective can be computed in linear time. We establish sufficient conditions for policy-optimality preservation and show that task and embodiment matching impose complementary constraints on the solution space. Experiments across domains with different state dimensions and morphologies show that GWτdIL consistently learns target-directed behavior and outperforms prior cross-domain imitation methods under comparable assumptions. Ablations isolate the roles of temporal ordering, directional information, and task–embodiment matching.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.