acceptodds
Under review as a conference paper at ICLR 2027

Scaling In-Context Online Learning Capability of LLMs via Cross-Episode Meta-RL

Abstract

Large language models (LLMs) achieve strong performance when all task-relevant information is available upfront, as in static prediction and instruction-following problems. However, many real-world decision-making tasks are inherently online: crucial information must be acquired through interaction, feedback is delayed, and effective behavior requires balancing information collection and exploitation over time. While in-context learning enables adaptation without weight updates, existing LLMs often struggle to reliably leverage in-context interaction experience in such settings. In this work, we show that this limitation can be addressed through training. We introduce ORBIT, a multi-task, multi-episode meta–reinforcement learning framework that trains LLMs to learn from interaction in context. After meta-training, Qwen3-8B, a relatively small open-source model, demonstrates substantially improved in-context online learning on entirely unseen environments: it matches or slightly exceeds GPT-4o while outperforming other baselines by a large margin on average. Beyond performance, we find that ORBIT exhibits adaptive exploration, revising its actions after failure without explicit reflection prompts. Scaling experiments further reveal consistent gains with model size, suggesting significant headroom for learn-at-inference-time decision-making agents.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.