acceptodds
Under review as a conference paper at ICLR 2027

Endogenous Curricula for Verifiable Self-Play: Projected Soft- Coverage–Information Certificates

Abstract

Self-play agents must adapt executable-task curricula without extinguishing valid program paths. We model this as KL-regularized control on a syntax-constrained prefix tree. The same projected soft- iterate yields a completed-program floor and an information-value certificate. A realized-row bound sharpens support control; direct performance identities, quadratic residual penalties, and contrast-aware estimator errors certify information advantage over an IG-ablated comparator without the full-objective optimal-policy oracle. A coherent finite Bayesian specialization gives expected entropy budgets, while identifiable likelihoods and accumulated row floors yield anytime posterior and finite-time identification bounds. Fixed-horizon two-time-scale tracking and epochwise expansion delimit the dynamical scope. In an executable Boolean-DSL study, a matched pure-information curriculum improves target success and semantic coverage over its reference. Shared-belief audits show conditional gains after KL costs, but the original residual-composed information bound is nonpositive throughout the tested configurations. The direct refinements have not been evaluated on those records; we separate mathematical guarantees, observed learning gains, and unresolved finite-time efficacy.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.