acceptodds
Under review as a conference paper at ICLR 2027

Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Memory Requirements

Abstract

Adapting long-horizon LLM agents to specific tasks requires efficient and flexible optimization under limited GPU resources. Although Reinforcement Learning (RL) has shown promise in single-turn LLM fine-tuning, this setting exposes two challenges: backpropagation makes full-parameter adaptation increasingly expensive for larger LLMs and longer contexts, while sparse rewards and branching interactions complicate long-horizon credit assignment. We argue that evolution strategies (ES) can often be a better choice for such \textbf{task-specific adaptation, offering 1) Model Scalability through forward-only full-parameter optimization, 2) Flexibility to integrate parameter updates with prompt-space optimization, and 3) Long-Horizon Scalability by avoiding explicit turn-level credit decomposition. Based on these insights, we propose Agentic ESOpt, a full-parameter ES framework supporting both train-time fine-tuning and on-the-fly adaptation during agentic test-time compute, together with a cosine perturbation schedule for balancing exploration and adaptation. Agentic ESOpt outperforms the strongest Agentic RL baseline by 13.19 percentage points in our longest-horizon Sudoku experiment, improves over Agentic GRPO on multi-turn tool-usage Math and DocVQA, enables full-parameter fine-tuning on a MoE LLM Qwen3-30B-A3B with 480GB GPUs, and improves existing automatic heuristic design pipelines in 28 of 36 comparisons. Code is available at https://anonymous.4open.science/status/Agentic-ESOpt-anonymous-197C.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.