acceptodds
Under review as a conference paper at ICLR 2027

NutriAgent-RL: Learning to Route Rule-grounded Agent Workflows for Personalized Clinical Nutrition Guidance

Abstract

Personalized clinical-nutrition guidance should adapt general dietary knowledge to each patient’s clinical context. For an AI system providing such guidance, the context determines whether it should provide dietary advice directly, add a caution, or recommend referral. We use clinical-nutrition redline rules to describe how patient conditions and risk triggers change the response to a dietary query. Each rule specifies which advice should be avoided and what response the system should provide. Applying such rules requires determining which evidence to retrieve and which checks to perform for each query. A common design applies the same large language model (LLM) agent workflow to every query, even when cases differ in their need for retrieval and checking. This design wastes computation on simple cases and leaves relevant constraints unchecked in more complex ones. Thus, workflow selection is central to redline-aware generation. NutriAgent-RL addresses this limitation by learning a query-dependent policy over complete agent workflows. Its contextual-bandit router uses parsed patient context to choose among four complete workflows: Direct Answer, retrieval-augmented generation (RAG), Multi-agent Reasoning, and Fixed Revision. During offline training, rule-grounded critics use applicable rule evidence to score the completed outputs of candidate workflows. These scores, together with penalties for over-refusal and inference cost, provide route-level rewards for offline policy learning of the router. To evaluate the learning of workflow selection, we further construct ND-Bench with 45,033 tasks generated from 449 canonical rules grounded in public nutrition resources and guideline-derived evidence. Compared with RAG, the strongest fixed workflow on Qwen2.5-7B-Instruct, NutriAgent-RL lowers the benchmark-defined escape rate by 7.1%, while using 49.8% fewer total tokens per task.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.