acceptodds
Under review as a conference paper at ICLR 2027

Learning to Teach, Learning to Play: Dual Reinforcement Learning with Active Trajectory Inquiry

Abstract

Foundation models can guide reinforcement learning (RL) agents, but useful guidance depends on understanding how a situation arose and where it can still be changed. In dynamic games, the information needed for this judgment may be scattered across a trajectory and become apparent only during investigation, making it difficult to specify a sufficient context in advance. We introduce TrajTeach, a dual reinforcement learning framework that trains a foundation-model teacher to investigate and improve an RL student's behavior. We develop TrajQL, a structured trajectory-query language through which the teacher locates relevant events, state changes, and objects, follows up on its findings, and decides whether, where, and how to intervene. Return gains over the student's continuation train the teacher's inquiry and intervention policy, while the resulting experience trains the student. Trajectory analyses across six games illustrate how inquiry connects observed problems to actionable corrections, including tracing delayed failures to earlier decisions, recovering missing task prerequisites, and revising action timing. Only the student is deployed, preserving its original inference cost.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.