Idol Learning: Probe Before Commit for Rare-Success Reinforcement Learning
Abstract
Rare successful trajectories can reflect either reproducible behavior or favorable environmental randomness. Directly reinforcing them risks consolidating lucky outcomes. We introduce Idol Learning, a probe-before-commit framework that treats a high-return trajectory as a behavioral hypothesis. After an ordinary learner update, Idol temporarily amplifies the trajectory's action likelihood, compares fresh reference and probe rollouts, and retains an evidence-dependent fraction of the same candidate displacement. This separates the strength of a diagnostic intervention from that of permanent learning. Under matched interaction budgets, controlled experiments show substantial gains with REINFORCE and earlier acquisition or improved final skill with PPO. Probe-strength sweeps distinguish useful temporary amplification from harmful unconditional commitment. However, stronger local masking reverses the gains: reward-based validation can accept updates that damage the intended skill. DeepSea shows earlier consolidation than successful-experience replay in a retrospectively selected discovery cohort; exploratory FrozenLake runs show higher mean online success with heterogeneous outcomes. These results support active validation of rare success when interventions expose informative behavioral differences, while identifying candidate construction, reward fidelity, and commitment calibration as limits to broader reliability.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.