VisEvolver: Hybrid Parameter-and-Memory Self-Evolving for Multimodal Agents
Abstract
Multimodal large language models are evolving from static visual reasoners into agents capable of active perception, tool use, and sustained interaction, yet how such agents can continuously acquire capabilities from their growing experience remains underexplored. Existing approaches largely follow two paradigms: external memory enables rapid and reversible reuse of past experience but remains dependent on inference-time context, while parametric adaptation can turn experience into intrinsic capability but risks internalizing instance-specific facts and incidental behaviors. We argue that a self-evolving multimodal agent should instead allow experience to progressively migrate from external memory into internal capability according to its reliability and transferability, which raises two fundamental challenges: **Causal Experience Attribution**, identifying which parts of multimodal experience truly contribute to successful decisions, and **Selective Experience Internalization**, determining which validated experiences should be consolidated into model parameters. To address these challenges, we introduce **VisEvolver**, a dual-timescale framework that organizes multimodal experience into a progressive lifecycle of **Experience → Evidence → Skill → Program → Weight**. VisEvolver extracts transferable experience through visual counterfactual attribution and success–failure trajectory comparison, and selectively internalizes repeatedly validated skills via on-policy privileged self-distillation while dynamically coordinating external memory and internal skill experts. Experiments show that VisEvolver not only exploits newly acquired experience for immediate adaptation, but also turns stable and transferable strategies into persistent model capabilities, providing a unified framework for multimodal agents to move from remembering experience toward learning from it.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.