Learning from Evolving Prompts: Adaptive Prompt Teaching for Reinforcement Learning
Abstract
Prompt optimization and reinforcement learning (RL) improve language models through different channels: one edits the text in the model's context, while the other updates its weights. Recent methods combine the two by alternately updating the prompt and model weights to improve performance. However, the benefit of an optimized prompt may disappear as RL updates the weights, and the prompt may eventually become redundant or even harmful. We propose Adaptive Prompt Teaching (APT), an RL framework that uses evolving prompts as temporary teachers during training and tests whether any prompt is worth keeping at deployment. APT first samples each training problem from its original, non-optimized input and resamples it under teacher prompts only when every sample fails and thus provides no learning signal. It then trains the model on the guided responses with the teacher part removed from the input, allowing the model without the optimized prompt to learn the same behavior. Teacher prompts are periodically rewritten based on the model's recent failures and retired once the model performs as well without them. Since some guidance may remain unabsorbed in the weights, a separate step searches for a prompt that still improves the model with its updated weights and uses it as the final deployment prompt for evaluation. On HoVer, WebShop, and SATBench with Qwen3-8B and Gemma3-12B, APT achieves the highest mean accuracy in all six settings, 1.3 to 4.1 points above the strongest baseline.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.