acceptodds
Under review as a conference paper at ICLR 2027

When Answers Agree but Reasoning Disagrees: Conflict-Aware Agentic RAG for Medical Reasoning

Abstract

Large language models (LLMs) show promise in medical question answering, yet extended reasoning can introduce factual errors and conflicting claims that undermine answer reliability. Recent multi-round retrieval-augmented generation (MA-RAG) mitigates these issues through dynamic retrieval using answer-level consistency signals, leading to state-of-the-art performance yet overlooking the conflicts within individual reasoning trajectories. Specifically, it will cost N times tokens to get the retrieval signal during answer sampling and fails when the model consistently produces incorrect answers through disordered reasoning. In this paper, we begin by analysing reasoning trajectories and explore how they can guide retrieval and repair direction. We observe that trajectories can be identified into three patterns: smooth, repeat and conflict. Notably, trajectories containing conflicting claims have substantially lower answer accuracy than repetitive or smooth reasoning, suggesting that these semantic patterns offer a useful signal for external refinement. Motivated by this finding, we propose a Conflict-aware Agentic RAG framework for medical REasoning (CARE), which allows for a more reliable reasoning process with fewer tokens by analyzing reasoning trajectories to guide targeted interventions. It is achieved by preserving smooth reasoning, truncating repetition, and retrieving evidence to resolve factual conflicts, then integrating relevant reasoning and evidence to produce more reliable answers. Experiments on eight medical benchmarks show that CARE improves average accuracy by 2% over the strong adaptive RAG baseline MA-RAG while reducing generated tokens by approximately 77%. On MedThink-Bench, CARE further improves alignment with expert reasoning by 12.81 points across four semantic metrics. These results demonstrate the potential of trajectory-level conflict detection for improving both the reliability and efficiency of medical reasoning. Our code is available at https://anonymous.4open.science/r/CARE-C3F1/.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.