MedSpaceAgent: Training-Free Tool-Grounded Spatial Reasoning for Medical Visual Question Answering
Abstract
Medical visual question answering (VQA) models often produce answers through implicit cross-modal reasoning, making it difficult to determine whether their predictions are supported by relevant image regions and valid anatomical relationships. Moreover, adapting these models to spatial reasoning typically requires task-specific fine-tuning and additional annotations. We introduce MedSpaceAgent, a training-free, tool-grounded agent for explicit and verifiable spatial reasoning in Medical VQA. Instead of directly generating an answer, MedSpaceAgent translates each question into an executable spatial program and dynamically invokes frozen medical perception tools for anatomical segmentation, abnormality localization, and region-level inspection. The resulting observations are organized into a spatial evidence graph, on which a deterministic executor computes anatomical, directional, topological, and distance-based relations. An evidence-aware verifier further assesses evidence completeness, geometric consistency, and agreement across tools, enabling the agent to replan its perception strategy or abstain when reliable evidence is unavailable. This design separates linguistic planning from spatial computation and provides an auditable reasoning trajectory without updating any foundation-model parameters. We evaluate MedSpaceAgent on spatially focused Medical VQA tasks, including compositional, counterfactual, and out-of-distribution settings. The proposed framework establishes a transparent alternative to end-to-end adaptation and provides a general foundation for reliable spatial reasoning in medical multimodal systems.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.