From Experience to Lessons: Organizing Visual Memory for Chest X-Ray Report Generation
Abstract
LLM agents have been found promising for completing complex tasks, where the episodes recording how an agent made decisions to obtain the outcomes can be maintained in a memory store for experience reuse to make better decisions. In this paper, we investigate the use of MLLM agents for chest X-ray report genera- tion. We first make use of a multimodal (vision-language) episodic memory store for maintaining the experience. We further propose to organize the episodes as a compact set of multimodal lessons which are expected to be more generalizable for reuse. We conduct experiments on MIMIC-CXR with 76 lessons learned for retrieval by the agent to enhance the report generation quality, resulting in im- provement on finding-level balanced accuracy by 2.90 points, more than retrieval over all 8,192 raw episodes (1.07 points). In addition, by allowing the agent to learn how to properly incorporate the retrieved lessons via rewarding corrected findings and penalizing those introducing errors, the gain over the agent without memory rises to 5.18 points and the introduced errors fall from 123 to 74. We also demonstrate that the lessons learned are transferable to different X-ray datasets, different agent frameworks, and different MLLM base models.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.