ESTA-Mem: Compiling Multi-Agent Trajectories into Auditable Memory for High-Resolution Visual Reasoning
Abstract
Cognitive science shows that human visual memory binds observations to their sources for reliable decisions. In recent years, multimodal large language models (MLLMs) have approached human visual reasoning ability, but processing high-resolution images remains challenging. Existing agent methods enhance the visual perception ability of MLLMs, and multi-agent systems further improve complex reasoning by splitting tasks and assigning them to specialized roles. However, as multi-agent interactions increase, the shared history grows longer without clear links between visual findings and reasoning conclusions. Inspired by cognitive science, we propose ESTA-Mem, an vent-ourced, yped, and uditable Memory system that integrates interaction history into an evolving task state. Linked structures explicitly bind candidate answers to corresponding observations and verification records, while a memory agent proposes updates. Each agent receives these relations through a view tailored to its role, avoiding repeated reconstruction from mixed histories. This clarifies the current task state and next steps for each agent, making memory an active driver of multi-agent visual reasoning. We test our method on five high-resolution benchmarks. ESTA-Mem improves the baseline by 2.6 and 4.2 percentage points on V and TreeBench, respectively, and outperforms the compared memory methods. Results show that memory organization itself matters, improving visual reasoning without stronger models or longer reasoning.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.