acceptodds
Under review as a conference paper at ICLR 2027

Test-Time Self Improvement by Scaling Unsupervised World Internalization

Abstract

Agents accumulate experience as they interact with an environment, but current systems typically reuse this experience through context or external memory while keeping the underlying model fixed. While flexible, this constrains the form of learning models can do especially under novel and complex environments. We ask whether weight updates can instead improve the model's ability to solve tasks by learning directly from its own experience, without feedback on whther its attempts succeed. We study world internalization, a training approach that converts information acquired through unsupervised exploration into parameter updates. World internalization trains on insights and summaries extracted from exploratory trajectories, together with a small fraction of trajectories that teach the model to integrate newly acquired knowledge with existing cognitive priors for reasoning and action. This enables the model to internalize environmental knowledge through continued training analogous to pretraining and midtraining while preserving its ability to use that knowledge for downstream tasks. We evaluate world internalization across three domains with distinct forms of environmental knowledge and interaction: Harvey's Legal Agent Benchmark, StudyBench, and Equational Theories. We observe that internalization reliably improves downstream task success despite receiving no supervision on improving task solving. Moreover, performance scales nearly log-linearly as internalization data increases from 10M to 100M tokens. In all, our results show that scaling the internalization of autonomous experience can enable self improvement.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.