EpisodeScope: Budget-Aware Evaluation of Diversity Corrections in Molecular Agents
Abstract
Reinforcement learning with verifiable rewards can reduce output diversity, and several corrections restore it within a rollout group or a selected batch. We ask whether local diversity improves discovery per oracle call when a multi-turn agent also controls its oracle spend. We introduce EpisodeScope, an evaluation protocol for multi-turn molecular agents under an unenforced oracle budget. It counts distinct molecules above a baseline-calibrated score threshold per charged call within a validation window and reports realised spend and coverage at common call counts. In four seed-matched Qwen3-4B pairs trained separately under the same per-step scoring cap, farthest-first selection raises within-episode fingerprint spread but lowers qualified discovery per call by – in every pair. Repeated evaluations concentrate on its first-position anchor. A registered, non-concurrent random-anchor control still trails uniform selection in all three same-machine pairs, and determinantal selection gives mixed results. On replayed candidate pools, avoiding already-scored molecules reduces repeats most consistently. Enabling the complete memory configuration raises discovery per call in all five prospectively specified pairs (median difference , one-sided sign test ), although one memory run spends calls per episode. At this operating point, higher local diversity need not improve qualified discovery per call.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.