MetaRule: Transactional Procedure Learning for Cross-Domain Harness Evolution
Abstract
Language-model agents rely on harnesses to organize context, memory, tools, and output handling, yet repeated execution does not explain when a local trace should become durable behavior. Existing model-centric approaches can store memories, revise prompts, or search harness configurations, but they lack a transaction for deciding when execution evidence is closed enough to persist, limiting its scope, and keeping learning separate from evaluation. We present , a transactional procedure-learning mechanism for self-evolving harnesses. It converts answer-free execution traces into typed, scoped, lifecycle-managed programs: replayable checkpoints expose telemetry to an isolated diagnoser; bounded same-thread revision and typed closure establish whether a procedure completed; and paired admission, provenance, and delayed publication control what a successor harness may execute. The compiler dispatches only trigger- and capability-matched operations in the next frozen state, while official verification and read-only held-out evaluation preserve the measurement boundary. Across scientific and competition-mathematics benchmarks, the complete treatment improves held-out performance over native harness evolution, with the clearest gains in registered cross-domain and cross-type transfer and in recovery from native negative transfer. Matched frozen-state and trajectory-shadow analyses place these gains at the complete harness-treatment level and quantify uncertainty around persistent procedure transfer. provides a replayable, auditable substrate for studying how execution evidence changes an agent's future behavior.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.