Decoupling Medical LLM Reasoning with AgentTool Reinforcement Learning
Abstract
Large language models (LLMs) have achieved strong performance on medical question answering, yet single-number accuracy obscures where they fail. We introduce a diagnostic evaluation framework that decomposes medical reasoning into three complementary abilities: , the possession of question-relevant medical knowledge; , the ability to retrieve such knowledge when answering a question; and , the ability to solve the question when the required knowledge is provided in context. We instantiate this framework on MedQA, MedMCQA, and MMLU-Med with dedicated probes for each ability. Evaluating six LLMs reveals that recall tracks downstream accuracy most consistently among the three measured abilities, whereas recognized knowledge alone shows weak and sometimes negative association with correctness. We further show that standard reinforcement learning (e.g., GRPO) improves primarily medium- and high-recall questions, leaving the low-recall regime largely untouched in our setting. Motivated by these findings, we propose , a reinforcement learning framework that addresses the low-recall bottleneck by making parametric recall an explicit, on-demand tool. The tool runs in an isolated context to generate relevant knowledge without disrupting final-answer reasoning. We further optimize reasoning and recall with decoupled group-wise objectives, using answer correctness for reasoning trajectories and a length-regularized recall-coverage reward for tool-generated recall. Across in-domain and out-of-distribution medical QA benchmarks, AgentTool RL outperforms standard GRPO and strong medical baselines on every dataset.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.