EMIT: Measuring the Reliability of Per-Prompt Decisions in RLVR Curricula
Abstract
Per-prompt controllers in reinforcement learning with verifiable rewards transform small groups of verifier outcomes into consequential training actions, yet the reproducibility and semantic fidelity of those decisions are rarely characterized. We introduce EMIT, a reliability-aware evaluation framework that treats adaptive curriculum control as a measurable decision process. EMIT formalizes exact-action reproducibility, connects difficulty concentration and decision boundaries to rollout requirements, and combines these predictions with source-traced implementation and verifier audits. Applied to frozen real-server rollouts, the framework reveals how sampling resolution and answer-format semantics govern controller behavior. A completion-token-matched multi-run study further shows substantial observed gains for both the released adaptive controller and a carefully tuned static curriculum over the no-curriculum reference. The adaptive controller attains the highest observed primary-task mean while closely tracking the tuned static curriculum across the broader suite. By making intermediate curriculum decisions auditable, EMIT provides a principled basis for reporting, sampling, pooling, and reliability-aware controller design.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.