acceptodds
Under review as a conference paper at ICLR 2027

State2State: Environment-Derived Mid-Training for LLM Agents

Abstract

Training LLM agents commonly relies on supervised fine-tuning from expert trajectories or online reinforcement learning over human-specified tasks with handcrafted verifiers. Though effective, both remain bottlenecked by externally specified tasks and supervision signals, limiting the scalability and diversity of agent training. We study an environment learning paradigm in which agents acquire interaction and manipulation capabilities solely through environment interaction, without externally specified tasks. We propose State2State, an environment-derived mid-training method that converts explored environment states into training objectives, challenging agents to reach a specified target state. By deriving tasks from environment exploration and verifying success through rule-based state matching, State2State provides scalable and verifiable training objectives without expert supervision or manual task design. Experiments across textual and mobile GUI environments demonstrate the effectiveness of State2State as a standalone environment-learning stage, with performance gains observed in both interaction modalities. As initialization for downstream RL, it further improves final performance and learning efficiency, retains a 3.5 to 6.6 point margin over RL given the same total number of update steps, and provides initial evidence of positive transfer across environments.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.