Continual Test-Time Learning with Evolution Strategies
Abstract
Large language models increasingly operate in open-ended environments, where they are expected to continually adapt to evolving tasks and distributions, often without access to external supervision. However, existing continual learning approaches often rely on external supervision, such as labeled data, demonstrations, or reward signals, which can be costly and unavailable in practice. In contrast, repeated test-time adaptation using self-generated supervision can lead to policy collapse, degraded reasoning, and catastrophic forgetting. In this paper, we propose **Evolution Strategy Reinforcement Learning (ESRL)**, a continual test-time learning framework designed to mitigate catastrophic forgetting during label-free continual adaptation. Specifically, ESRL constructs a population of locally perturbed policies around the current model to explore multiple neighboring parameter regions throughout adaptation, and then samples trajectories from those regions. Unlike trajectory-space exploration from a single evolving policy, this parameter-space population maintains diverse sources of self-generated supervision even when the central policy begins to drift. The resulting trajectories are aggregated into label-free group-relative learning signals to update the central student. Through extensive experiments on single-task test-time adaptation and multi-task continual learning across mathematical reasoning, general knowledge, and medical domains, ESRL achieves relative gains of up to 31.7% over the base model and 16.6% over the second-best method. Notably, ESRL substantially reduces catastrophic forgetting and maintains the strongest final performance at the end of continual learning.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.