acceptodds
Under review as a conference paper at ICLR 2027

MART: Multi-Agent Hyperparameter Tuning for LLM Reinforcement Learning

Abstract

LLM reinforcement learning (RL) is computationally intensive, and well-chosen hyperparameters can improve both training efficiency and model performance. However, parameter tuning is highly dependent on expert knowledge and repeated trial-and-error experiments, incurring considerable human and computational overhead. Recent agent-based approaches have demonstrated the potential of automating system configuration optimization for LLM training and serving. However, extending such approaches to RL requires accounting for its distinct execution and learning dynamics: rollout and policy-update phases impose different memory and compute demands, while response lengths evolve during training, changing the workload and complicating configuration evaluation. Moreover, effective tuning must consider both execution efficiency and the effects of learning hyperparameters on training stability and task performance. To address these challenges, we propose MART, a multi-agent framework with human-in-the-loop review for hyperparameter tuning in RL post-training. Guided by phase-wise resource measurements and training feedback, specialized agents propose and assess candidate configurations. A human reviewer approves each selected candidate before execution or rejects it with feedback for subsequent revision. The search proceeds in two stages: short-horizon trials identify system configurations that improve end-to-end throughput while retaining a prescribed GPU memory margin; longer-horizon trials then optimize learning hyperparameters with the system configuration fixed. During learning tuning, online training-health assessment supports early termination of trials exhibiting persistent anomalies. Experiments on GRPO-based mathematical reasoning show that, relative to baseline configurations, MART improves throughput averaged over updates 3–5 by 13.5% and 7.7% for Qwen3-8B, and by 69.5% and 91.3% for Qwen3-0.6B-Base, under single-/dual-node deployments, respectively. Across AMC23, OlympiadBench, AIME24, and AIME25, MART improves mean pass@1 by 1.98–2.51 percentage points for Qwen3-8B and 0.74–0.81 percentage points for Qwen3-0.6B-Base.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.