acceptodds
Under review as a conference paper at ICLR 2027

How Should LLMs Teach RL Agents? Action Assistance and Competence-Adaptive Release

Abstract

Reinforcement learning (RL) in complex decision-making tasks often requires extensive environment interaction. Large language models (LLMs) can provide state-conditioned action advice, but success with teacher assistance does not imply that an RL agent can solve the task independently. We propose a competence-adaptive framework for LLM-guided RL that distinguishes two roles of teacher guidance: action intervention, which changes the trajectories collected from the environment, and demonstration supervision, which supports RL policy updates. During training, the framework executes the LLM teacher’s action with an adaptive probability. State–action pairs from actual interventions are stored and replayed to provide supervision alongside the RL objective. A fuzzy controller adjusts the intervention probability based on policy uncertainty, teacher-free evaluation success, and agreement between the RL agent’s proposed actions and the teacher’s actions, while retaining a positive base weight for demonstration supervision. We evaluate the framework on three MiniGrid task families, DoorKey, UnlockPickup, and MultiRoom, and Env2. Across six principal MiniGrid configurations, median final teacher-free success ranges from 90% to 100%. On DoorKey 5×5, the median number of environment steps to first reach 95% teacher-free success decreases from 64,000 to 6,000. Ablations show that persistent teacher-action takeover can impair independent performance, and that removing demonstration supervision slows learning on some tasks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.