acceptodds
Under review as a conference paper at ICLR 2027

Manifold Bandits: Bayesian Curriculum Learning over the Latent Geometry of Large Language Models

Abstract

Reinforcement learning (RL) is a central approach for improving reasoning capabilities in large language models (LLMs), where training efficiency depends critically on how problems are sampled during optimization. Existing adaptive curriculum learning methods typically prioritize prompts of intermediate difficulty, treating problem selection as a standard bandit problem with independent arms and overlooking the structured, heterogeneous nature of the task space. In this work, we instead frame problem sampling as a manifold-structured bandit problem with endogenous non-stationarity, where relationships between tasks and evolving model capabilities jointly influence learning dynamics. To operationalize this perspective, we introduce **Bayesian Manifold Curriculum (BMC)**, a structure-aware framework that leverages latent representations of LLMs to organize problems into a hierarchical task tree and applies Bayesian learning to guide sampling. Across multiple training setups and evaluation benchmarks, we find that different sampling strategies induce non-trivial tradeoffs between productivity (learning signal), diversity (coverage of the task manifold), and utility (evaluation relevance). These results show that prioritizing difficulty alone is insufficient to ensure strong downstream performance and highlight the importance of incorporating structure and diversity in problem sampling.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.