acceptodds
Under review as a conference paper at ICLR 2027

Real-Time Heuristic Learning: Compiling Online Evidence into Executable Policies

Abstract

Real-time agents must act promptly, but handling changing states does not establish adaptation to changing rules. We introduce Evolving Freeway and Evolving Snake, extending Real-Time Reasoning Gym with repeated hidden rule changes within continuing episodes. Phase boundaries are public and active rules hidden. Transitions provide evidence, and world state persists. Reactive, Planning, and AgileThinker retain discoveries as prompt evidence, action sequences, or textual guidance rather than installed policy revisions. We propose Real-Time Heuristic Learning (RTHL), an evidence-to-code framework compiling transition evidence into reusable executable rules. A deterministic heuristic selects actions. Post-transition diagnosis drives bounded code revisions, screened by program checks and replay, with phase scoping and rollback governing reuse. Across 576 episodes, RTHL leads all six pooled comparisons of mean terminal scores across environments and budgets, gaining 8.5–12.7% over Reactive on Freeway and 43.4–49.4% on Snake. At 1K, updating the same fixed initial controller adds 5.00 Freeway points and 10.25 Snake points. RTHL leads 12 of 14 Hard comparisons across seven budgets (1K–32K), including seven on Snake. RTHL enables reusable adaptation without weight updates.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.