acceptodds
Under review as a conference paper at ICLR 2027

MetaClaw: Just Talk – An Agent That Meta-Learns and Evolves in the Wild

Abstract

Large Language Model (LLM) agents deployed in dynamic environments often remain static, becoming stale as task distributions drift. Existing solutions either store raw trajectories without knowledge distillation, maintain skill libraries disconnected from weight optimization, or require service downtime for retraining. We present MetaClaw, a continual meta-learning framework that jointly evolves a base LLM policy and a skill library of reusable behavioral instructions through two complementary mechanisms. First, skill-driven fast adaptation synthesizes new skills from failure trajectories via an LLM evolver, providing immediate, zero-downtime improvements. Second, opportunistic policy optimization performs gradient-based updates via cloud-based RL during user-inactive windows, managed by a scheduler (OMLS) that monitors activity signals like Google Calendar. To ensure data integrity, a versioning mechanism strictly separates support data (for skill evolution) from query data (for RL updates), preventing stale reward contamination. Built on a proxy-based architecture, MetaClaw scales to production LLMs without local GPUs. Experiments on MetaClaw-Bench and the AutoResearchClaw pipeline demonstrate that skill-driven adaptation alone improves accuracy by up to 32% relative. The full framework elevates Kimi-K2.5 from 21.4% to 40.6% accuracy, nearly matching the GPT-5.2 baseline (41.1%), while achieving an 8.25 gain in end-to-end task completion and an 18.3% boost in robustness.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.