Not Every Workflow Helps: Discover, Verify, and Assemble Workflows for Agentic RAG
Abstract
Agentic retrieval-augmented generation (RAG) uses workflows to retrieve and process evidence. However, training RAG policies is computationally expensive, while handcrafting workflows requires repeated case review and validation. Reusing past workflows offers an alternative, but previously effective workflows may become counterproductive under new conditions. We investigate failure modes in RAG workflow reuse and identify three phenomena: Workflow Reversal, Success Mismatch, and Noise Amplification Loop. These findings motivate verifying whether a workflow provides a benefit under the evidence conditions in which it is applied. To avoid ineffective reuse, we propose REMEDY-RAG with three stages: Discover, Verify, and Assemble. It extracts workflows from past executions, learns evidence-conditioned assignments from paired gains over no intervention, and independently verifies these assignments. An external workflow map selects a verified workflow or no intervention based on the initial evidence state. With the same verified workflow, evidence-conditioned admission raises macro F1 from 47.0 to 50.3 across three QA benchmarks. Across five benchmarks, REMEDY-RAG with a frozen Qwen3-8B executor reaches 50.8 macro F1, surpassing the strongest baseline while using 80.1% less offline adaptation compute. The map complements policy training and improves the performance of Qwen3-32B and GPT-5.4 without target-side rediscovery, reverification, or tuning. These results highlight the value of learning and verifying when past experience improves agentic RAG.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.