UnCurl: Reliability-Conditioned Partial Sharing for Joint Curriculum and Exploration Control
Abstract
Efficiently training large language model (LLM) agents requires deciding both which tasks to practice and how much to explore within each task. Existing controllers usually estimate these decisions independently, although both can depend on shared capabilities; forcing one shared uncertainty representation, however, can propagate proxy error and cause negative transfer. Our contributions are threefold. We formulate curriculum–exploration coordination as conditional partial sharing between task-level learning progress (LP) and state-action information gain (IG). We introduce UnCurl, which retains private LP and IG factors and admits a shared predictive factor only when calibrated dependence, reliability, held-out utility, effective sample size, and drift tests pass. A monotone soft gate controls sharing strength, while an independently trained shadow posterior provides fail-closed fallback. On ALFWorld, UnCurl reaches 91% unseen success versus 86% for interaction-matched Independent LP+IG, raises normalized AULC from 73 to 81, and reduces interactions to the 80% development threshold from 43k to 30k. It lowers harmed-task fraction from 0.48 under Full-share LP+IG to 0.16 and reaches 44% success on WebArena-Lite. UnCurl uses 37.7% more GPU-hours than Independent LP+IG.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.