acceptodds
Under review as a conference paper at ICLR 2027

LongRx: Closed-Loop Medication Recommendation with Action-Coverage Reinforcement Learning

Abstract

Medication recommendation is commonly formulated as one-shot set prediction, although prescription construction naturally involves iterative information acquisition, medication adjustment, and drug–drug interaction verification. We propose LongRx, a closed-loop medication decision agent that sequentially constructs prescriptions through five parameterized actions: RETRIEVE, ADD, DELETE, VERIFY, and EXIT. Starting from an empty candidate set and prescription, LongRx learns state-dependent decision sequences rather than following a predefined workflow. To improve sequential learning, we combine terminal prescription supervision with process-level feedback on local prescription changes. Our Coverage-Aware Tree Optimization (CATO) combines opportunity-aware action coverage with tree-relative credit, broadening exploration of under-explored valid decisions and comparing their downstream outcomes from the same prescription state, while concrete medication edits and retrieval queries remain policy-generated. In an idealized fixed-state, single-draw model, the coverage correction contracts pairwise action log-odds. Experiments on MIMIC-III, MIMIC-IV, and eICU show improved overlap with recorded prescriptions and lower rates of DrugBank-listed interacting medication pairs relative to RES-MR. These metrics characterize agreement with historical prescribing and a database-based interaction proxy; they do not establish clinical appropriateness or improved patient outcomes.Our code is available at https://anonymous.4open.science/r/LongRx.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.