acceptodds
Under review as a conference paper at ICLR 2027

SEAL: Synergistic Co-Evolution of Agents and Learning Environments

Abstract

Large Language Model (LLM) agents increasingly learn through interaction, yet their training process often evolves unevenly: the policy changes with experience while the learning environment continues to provide largely fixed forms of support. As training progresses, this creates a growing mismatch between what the agent currently struggles with and what its learning process emphasizes. We call this phenomenon Agent–Learning Environment Misalignment (ALEM). We introduce SEAL, a framework for synergistic co-evolution of agents and learning environments. SEAL turns verifier-grounded execution failures into a shared diagnostic signal that coordinates two complementary forms of adaptation: it reshapes the support available in subsequent interactions and adjusts how the policy learns from collected trajectories. In this way, the learning environment follows the agent’s evolving failure profile while both sides remain grounded in the same executable feedback. Across three backbones on BFCL V3, SEAL consistently improves over matched Vanilla RL, reaching 40.25% versus 30.75% with Qwen2.5-7B. A matched factorial ablation further yields a +5.00-point interaction estimate for the coupled configuration. The same trained checkpoints also show positive transfer to the evaluated BFCL V4 and -bench domains. These results suggest that coordinating how an agent learns with the environment it learns from is a promising direction for self-evolving tool-use agents.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.