acceptodds
Under review as a conference paper at ICLR 2027

Bilevel Incentive Optimization for Language Model Self-Improvement without Human Labels

Abstract

Self-improvement enables large language models to learn from self-generated problems and automated feedback, but its effectiveness depends on the incentives governing problem generation and solution learning. We propose a bilevel post-training paradigm that jointly adapts these incentives and the policies they induce without human-provided solution labels. At the lower level, a questioner and a solver interact in a general-sum game under role-specific rewards. At the upper level, a shared incentive model minimizes the equilibrium solver's expected task failure. We propose BiGAP, a bilevel game-aware penalty method that constructs policy-specific objectives from Nikaido–Isoda gap functions and aggregates both players' gaps for incentive learning. Its training protocol coordinates question generation, separate practice and target-response sampling, policy optimization, and incentive adaptation, with optional target-question refresh. For fixed distributions and a fixed penalty parameter, we show under suitable regularity assumptions that, after exact-oracle iterations, the average squared operator residual is . The analysis further relates optimization accuracy and penalty strength to approximate Nash equilibrium guarantees. Across nine mathematical reasoning and formal theorem-proving benchmarks, BiGAP consistently outperforms fixed-incentive self-improvement. These gains are achieved without human-provided solution labels during self-improvement.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.