acceptodds
Under review as a conference paper at ICLR 2027

EHRGym: A Diagnostic Benchmark for Offline Reinforcement Learning

Abstract

A low offline-RL score on clinical-style logs can reflect little room to improve, missing information, or failed learning. EHRGym separates these possibilities in one synthetic ventilator task with hand-specified responses and rewards. Its reference ladder runs from the exactly computed best constant setting to a hidden-state oracle. Planners that know the simulator probe what past settings, timing, and measurements can support. A clock-only plan follows a fixed schedule of settings indexed by time and uses no measurements. On varied logs, this plan reaches about a quarter of the way from floor to ceiling. DQN/FQI and MedDreamer see the full record yet stay near the floor. With hidden-state access, DQN/FQI reaches more than halfway, exposing difficulty learning from the partial record. Under repetitive logging, CQL and Decision Transformer often copy starting settings and score below their stochastic behavior policy. Researchers can use information-matched references to interpret low scores and inspect chosen actions for logging habits.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.