acceptodds
Under review as a conference paper at ICLR 2027

SimpleHRL: Minimal Hierarchical Structure Improves Out-of-Distribution Generalization in RL Post-Training of LLMs

Abstract

Reinforcement learning (RL) post-training has become a central paradigm for improving LLM reasoning, yet how the structure of the reasoning trace during training affects out-of-distribution (OOD) transfer remains poorly understood. Hierarchical RL (HRL) offers a natural inductive bias-separating what to do from how to do it-intended to encourage abstract, transferable reasoning, but HRL has been studied almost entirely outside the LLM RL post-training regime, leaving open whether its structural prior alone is enough to deliver OOD gains in LLMs. We introduce SimpleHRL, a deliberately minimal, HRL-inspired structural prior for LLM RL post-training: it imposes a hierarchical plan/execution structure on the reasoning trace during training, with no separate critic, no multi-level reward shaping, and no objectives beyond what flat baselines already use. Despite this minimality, when trained only on mathematics, SimpleHRL improves average OOD Pass@32 by  pp and Pass@1 by  pp over GRPO+CoT and entropy-exploration baselines across a broad OOD benchmark. To understand this effect, we conduct a series of mechanistic analyses and reveal two complementary insights into how hierarchical structure helps: SimpleHRL's plan-level reasoning becomes markedly more task-agnostic, supporting broader and more transferable strategy coverage; and hierarchical training concentrates gradient updates on the pivotal decision tokens that most determine outcome quality. Together, these findings argue for treating the structure of the reasoning trace during training, beyond surface-level CoT formatting, as a productive lever for OOD generalization in LLM post-training.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.