CausalMem: Counterfactually Validated Memory for Training-Free Visual Reasoning Agents
Abstract
Autonomous visual agents should improve as they accumulate experience, yet updating large vision-language models after every interaction is impractical. External memory offers a training-free route to self-evolution, but naive memory accumulation can amplify visual shortcuts: an agent may remember background, style, layout, or co-occurring objects rather than the evidence that actually determines the answer. We introduce , a training-free memory framework that writes visual memories only after counterfactual validation. CausalMem alternates between wake-time reasoning and sleep-time consolidation. During wake, a frozen VLM solves new image reasoning tasks by retrieving validated semantic rules, hard negatives, and anchor episodes. During sleep, the agent proposes causal and nuisance visual factors, edits stored images as noisy interventions, and re-answers the edited cases. A rule is promoted to long-term memory only if the answer is invariant to nuisance edits and sensitive to edits of the hypothesized causal evidence. We formalize this process as noisy interventional memory validation and derive a finite-sample bound on the false acceptance of spurious rules. Across seven visual reasoning benchmarks, CausalMem improves Qwen2.5-VL-7B and Qwen3-VL-8B by 4.18 and 3.29 average points over vanilla inference without updating model parameters. Component ablations, stream-order and inductive-split controls, cross-benchmark transfer, and experiments with other editor and VLM families support that, on the tested settings, the gains come from counterfactually validated memory rather than passive retrieval, suggesting that self-evolving visual agents must learn not only how to recall experience, but which rules deserve to be remembered.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.