acceptodds
Under review as a conference paper at ICLR 2027

MIMUS: Post-Hoc Refinement of Frozen Rewards via Residual Evolution

Abstract

Reinforcement learning systems often rely on rewards that are costly to redesign and may not fully capture task objectives. Inspired by inheritance and variation in natural evolution, we introduce MIMUS, a post-hoc reward refinement framework that freezes an existing base reward and evolves a compact, low-rank additive neural residual. The residual is initialized to output zero, ensuring that evolution starts exactly from the base reward. Before evolution, a Signal Gate checks whether candidate reward comparisons reveal a reliable optimization signal amid policy-training noise. The residual is subsequently evolved using Covariance Matrix Adaptation Evolution Strategy (CMA-ES), guided by teaching fitness computed from the native returns or task-aligned metrics of policies trained under candidate rewards. This approach enables reward optimization without residual supervision or differentiation through policy training. Beyond the evolutionary search, long-horizon audits guide reward selection for final policy training. We evaluate MIMUS on four tasks in Gymnax and MJX with handcrafted, LLM-generated, and learned neural base rewards. Across these tasks and reward sources, MIMUS improves downstream policy performance while leaving the base rewards unchanged, demonstrating that residual evolution offers a practical complement to reward design.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.