acceptodds
Under review as a conference paper at ICLR 2027

Remember Right, Look Ahead: History Notes with Future Prediction for LLM Agents

Abstract

Language-model agents trained with multi-turn reinforcement learning (RL) often see only their last few steps, so they lose what they observed earlier. An agent can keep this information in a short text note that it writes at every step and reads at the next, but RL rewards the agent for completing the task and never checks whether the note is correct. We introduce History Notes with Future Prediction (ForeNote), which supervises the note directly. In the text environments we study, everything an agent learns during an episode comes from text feedback, such as a confirmed pickup or an opened drawer, so a small rule-based tracker can read the same feedback and write the correct note, a reference note. In both supervised fine-tuning and RL, ForeNote uses this reference as two training signals: it trains the agent to write its current note correctly (remember right) and to predict the next note after each action (look ahead). At test time, the agent writes and reads its own note without the tracker. With Qwen2.5-3B-Instruct, ForeNote reaches 94.6% and 95.2% success on seen and unseen ALFWorld scenes and solves 59.1% of ScienceWorld tasks. Our analysis on ALFWorld shows where these gains come from. Without a note, an agent is twice as likely to revisit a searched place just after it leaves its five-step window, whereas the note of ForeNote still lists 99.6% of such places, and the agent relies on it: erasing the note at test time costs 8.5 points of unseen success. Without the two signals during RL, unseen success falls by 20.8 points, and giving the agent a correct note at test time does not bring it back. Unseen success also falls by 4.8 points without next-note prediction in RL, and by 3.9 and 15.7 points when the agent instead predicts the next observation or another game's next note.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.