Learning to Learn Experiential Knowledge
Abstract
Language agents can solve individual tasks effectively, yet they remain limited at accumulating experiential knowledge: converting prior task trajectories into reusable knowledge in free-form text that improves performance on future tasks. Prior work prompts strong models to write rules, skills, and strategies into an in-context notebook, but few approaches train models to improve this capability over long task streams. Training is challenging because the utility of each update is only revealed through later tasks. We study experiential knowledge accumulation through a persistent in-context notebook that is read by a task model and rewritten after each task by a notebook model. We propose NoteTaker, which trains the notebook model in two stages: SFT on teacher-generated notebook updates, followed by RL on short task streams, where each update receives credit only from subsequent task rewards that it could influence. Across three benchmarks, separately trained notebook models improve task models when deployed in online 40-task streams. These improvements transfer to held-out tasks, including AppWorld Test-Challenge tasks that involve applications and APIs absent from notebook model training. On AppWorld Test-Challenge, our finetuned Qwen3.5-9B notebook model increases task success from 0.0% to 25.8%, compared with 23.0% when Fable-5.1 serves as the notebook model. Analysis shows that SFT establishes basic experience abstraction and knowledge organization, while RL aligns notebook updates with downstream task performance, consistently improving over SFT across benchmarks and model sizes.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.