Benchmarking In-Context Experiential Learning across Heterogeneous Action and Feedback Channels
Abstract
Real-world agents regularly contend with a sequence of tasks under a shared partially observable environment. Partial observability requires the agent to elicit feedback strategically, synthesize observations into actionable experience, and thereby refine its actions; shared environment requires the agent to not myopically act for the task at hand, but also for the subsequent tasks that may benefit from accrued experience. We refer to this adaptive cross-episode learning as experiential learning. Recent agent evaluations have begun to measure experiential learning in conversational or coding settings, where interactions occur only over a single channel. In contrast, agents in the wild tackle far more complicated scenarios, where experiences arrive through multiple senses and various distinct modes of actions are available. Web-based product recommendation is one such rich setting: a recommender agent must learn a customer's preferences over many shopping episodes, acting through natural-language dialogue and dynamic on-site product displays, while the customer reveals their preferences both in their replies and in logged browsing behaviors. Note that the latter signal is far noisier and frequent than text, so a performant agent must infer preferences by parsing the abundant yet opaque logs alongside informative yet sparse user responses. Evaluating state-of-the-art LLM agents in this setting, we found that agents can leverage multiple interaction channels to learn within a task; yet show no improvement across tasks. Decomposing experiential learning into synthesis, belief, and action, we identify three failures modes in the agents: (1) difficulty extracting useful information from the noisier channel, (2) forming low quality beliefs, both about the current customer modeling and the informativeness of its actions, and (3) failure to explore in proportion to its uncertainty. Our results suggest closing this gap toward robust experiential learning will require agents that overcome these failure modes.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.