Self-Configuring Hierarchies with Meta- Adaptive Hierarchical Reinforcement Learning for Multi-Agent Cooperative Control
Abstract
Hierarchical reinforcement learning organises cooperative multi-agent behaviour into levels of temporal abstraction, but the structural parameters of the hierarchy (temporal commitment intervals, per-level discount factors, and inter-level feedback gains) are almost always fixed by hand before training. We call this the selfconfiguration gap: the hierarchy cannot adapt its own structure to the task, so every new environment requires a fresh hyperparameter search, and the resulting architecture is misaligned whenever the task horizon or team size departs from the calibration setting. We introduce Meta-MOHRL, a bilevel framework in which a meta-controller maps a compact, observable task descriptor to the full structural configuration of a three-level multi-objective hierarchy once per episode; the inner loop trains the hierarchical policy under that configuration and the outer loop updates the meta-controller from observed returns across a distribution of tasks. What is meta-learned is not a policy initialisation or a set of policy weights but the temporal structure of the induced semi-Markov decision problem itself: the commitment intervals set the effective horizon each level must reason over, and the per-level discounts set the timescale at which each objective is optimised. We give a scaling argument showing that balancing estimation cost and commitment regret yields an optimal commitment interval that scales with the square root of the horizon. On a cooperative SUMO-RL traffic-signal benchmark with a heterogeneous vehicle fleet and 15 randomised network topologies, Meta-MOHRL outperforms fixed-structure hierarchical and flat multi-objective baselines in total reward and Pareto hypervolume, and recovers a per-level discount structure that was never imposed as a design choice. The meta-controller executes once per episode and adds negligible test-time overhead.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.