acceptodds
Under review as a conference paper at ICLR 2027

RICE: Response-Informed Context Evolution for Agent Learning

Abstract

Interactive agents increasingly summarize completed trajectories as natural-language skills and reuse them across policy updates. This persistence allows past experience to support later decisions, but it also creates a mismatch: stored skill content can remain available while the policy that interprets it continues to change. Task outcomes evaluate the quality of sampled behavior, yet they do not record how retrieved context changes the current policy's predictions. On ALFWorld/Qwen3-1.7B, of the recorded response histories associated with stored skills have tied outcome differences; within this subset, the 95th percentile of the response magnitude is times the 5th percentile. This observation shows that outcomes and policy responses describe different properties of the same retrieval event. We define the local policy response as the difference between next-token distributions with and without retrieved context at the same decision prefix. We then introduce Response-Informed Context Evolution (), which accumulates these responses across retrieval events and associates the resulting history with the participating skills. Returns select which observed behavior provides supervision, response history determines how strongly accepted examples contribute, and response history together with past outcome evidence controls which skills remain available for future retrieval. Paused skills remain stored and can return to the available set when their outcome utility recovers. Across three model sizes on ALFWorld, WebShop, and Search-QA, achieves the highest listed overall success on ALFWorld and exact success on WebShop, while the Search-QA comparisons are mixed. In a controlled ALFWorld/Qwen3-1.7B comparison, validation success increases from for the Base configuration to with response weighting and with response-informed retrieval control. These results support treating retrieval as both context provision and feedback about how stored skills affect an evolving policy.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.