acceptodds
Under review as a conference paper at ICLR 2027

MuSEAgent: A Multimodal Reasoning Agent with Stateful Experiences

Abstract

Multimodal agents have shown growing ability to interleave tool use with multimodal observations for complex visual reasoning, yet effectively reusing experiences from past multimodal interactions remains challenging. Existing memory paradigms struggle in complex multimodal environments: coarse trajectories flood the context window with visual redundancy while obscuring fine-grained errors at intermediate steps; ungrounded scalar feedback offers no procedural guidance for visual tool execution; and holistic representations of multimodal states conflate visual appearance with tool histories, misguiding experience retrieval. To resolve these dilemmas, we introduce MuSEAgent, a multimodal reasoning agent with stateful experiences that reframes multimodal experience reuse across three foundational dimensions: (1) granularity, decomposing multimodal interaction trajectories into atomic state-action transitions anchored to intermediate decision states; (2) supervision, synthesizing outcome-conditioned natural-language visual tactics via hindsight reasoning to provide explicit procedural rules; and (3) representation, organizing heterogeneous multimodal states into a compositional multi-view space, queried via an iterative Deep-and-Wide search to dynamically align visual observations, task semantics, and tool execution histories. Extensive evaluations across four demanding multimodal benchmarks demonstrate that MuSEAgent consistently elevates diverse base models across parameter scales, surpassing trajectory-level baselines by up to 7.97 points and outperforming state-level scalar-value alternatives by 10.42 to 12.00 points. Remarkably, stateful experiences transfer robustly to unseen benchmarks without in-domain interaction, outperforming matched trajectory-level transfer by 7.8 points.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.