3DMedEvo: Self-Evolving VLM Agents with Evidence-Grounded Multimodal Skills for 3D Medical Image Analysis
Abstract
3D CT analysis integrates volumetric perception with clinical reasoning, and improves as experience accumulates over successive readings. Existing 3D medical methods typically use task-specific models or task-agnostic end-to-end predictors and rely on external teacher supervision. Memory-augmented approaches instead reuse experience across cases but often retain raw, noisy, and poorly curated interaction traces as text that does not fully preserve spatial detail. More importantly, they lack adoption-aware estimates of memory utility for future reasoning. Here, we present **3DMedEvo**, a **self-evolving runtime framework** that distills informative trajectories from frozen vision–language model (VLM) agents reading CT scans into structured multimodal skills through outcome-guided reflection, without additional teacher supervision. The resulting skills combine reusable procedures with visual exemplars in a repository organized into multiple branches, with guidance grounded in clinical evidence. 3DMedEvo further calibrates skill reliability from observed utility using an adoption-aware **Skill Credit Score (SCS)**, which guides retrieval for each task and uncertainty-aware governance. In turn, retrieved skills support reasoning by frozen VLM agents within a closed loop of skill acquisition, evaluation, and revision. Experiments on two established clinical benchmarks covering thoracic and abdominal CT show that 3DMedEvo outperforms reasoning agents using tools and representative agents using memory across recognition, quantitative measurement, visual reasoning, and medical reasoning tasks. The framework also generalizes across VLM backbones, offering a practical approach to self-evolution in 3D medical image analysis. All code and data will be made publicly available.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.