acceptodds
Under review as a conference paper at ICLR 2027

ATLAS: From World Laws to Skills in LLM-Guided Hierarchical Reinforcement Learning

Abstract

Long-horizon reinforcement learning (RL) often relies on structured skill specifications to overcome sparse rewards and long dependency chains. Existing LLM-Hierarchical Reinforcement Learning (LLM-HRL) methods typically ask language models to both infer environment dynamics and derive skill preconditions, effects, and dependencies, making specification errors difficult to avoid. We propose ATLAS (Acquiring and Testing Laws for Actionable Skills), a world-law-driven LLM-HRL framework in which explicit world laws connect language-model reasoning with skill construction and learning. ATLAS represents environment dynamics as a BC+ world model and deterministically derives executable skill specifications from accepted laws for hierarchical RL. Test-driven development provides validation constraints on law construction and revision through symbolic unit tests and environment-level tests. To acquire and revise laws from experience, ATLAS uses successful and failed trajectories together with counterfactual experiments to check candidate world laws. Accepted revisions propagate to skill specifications and training, translating evolving environment knowledge into continually updated execution capabilities. With detailed Craftax-Classic materials, the generated laws support operator grounding and achievement completion across resource and tool dependencies; in an unfamiliar Mars world, interaction evidence supports law acquisition without supplied environment-specific rules.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.