Mentor-Baseline Implicit Imitation Learning from a Single Action-Free Mentor Trajectory
Abstract
In online reinforcement learning (RL), we investigate whether a learning agent (observer) can benefit from experience acquired by another agent (mentor) in addition to its own interaction with the environment. The mentor need not be an expert, the two agents may have different state and action spaces, and the observer is given only a single action-free mentor trajectory of states and rewards. In this setting, we propose Mentor-Baseline Implicit Imitation Learning (\mbiil), which uses the mentor's return-to-go and approximate mentor–observer state correspondence to form state baselines and propagates them to guide the observer's action-value learning. We further formalize mentor-derived baseline optimism, which can drive exploration but can also bias value learning, and bound the resulting policy performance loss. In sparse-reward environments, we evaluate \mbiil using a simple ball as the mentor and observers with substantially different embodiment and control requirements, including a high-degree-of-freedom Ant for maze navigation and a Hopper for obstacle crossing. Our results show that a single action-free mentor trajectory can facilitate learning across substantial embodiment differences, with \mbiil being the only compared method to learn to reach the goal in AntMaze. They also show that the observer can achieve a higher return than the mentor, while mentor–observer correspondence too coarse to distinguish states of different values can induce overly optimistic baselines and impede learning, consistent with our theoretical analysis.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.