acceptodds
Under review as a conference paper at ICLR 2027

TextualRL: Textual Reinforcement Learning for Context Management

Abstract

Textual context optimization allows LLM agents to adapt to task feedback without changing their parameters. However, outcome labels alone do not identify which behaviors to preserve or correct. Instructions inferred from individual trajectories can also overgeneralize across tasks, excluding behavior observed in successful executions under different conditions. To address these problems, we introduce TextualRL, a textual reinforcement learning framework that combines trajectory comparison with evidence-based refinement of context updates. Multiple Trajectory Analysis contrasts successful and failed trajectories within a task. Across tasks whose trajectories all succeed or all fail, it identifies behaviors worth preserving or formulates repair hypotheses, respectively. Parallel analysts produce proposed edits and evidence cards. Cross Group Refinement reviews these edits against relevant cards pooled across groups, retaining, narrowing, or revising instructions based on supporting observations and exceptions. Reviewed edits form a textual patch, and held-out validation determines whether to adopt the resulting context. We evaluate TextualRL with five target agents on six benchmarks. Averaged over the six benchmarks, TextualRL improves test accuracy over the no-skill baseline by 24.4 and 15.3 percentage points for GPT-5.5 and Qwen3.8-27B.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.