acceptodds
Under review as a conference paper at ICLR 2027

LEAH: Jointly Training LLM Agents to Act and Revise Their Harness

Abstract

Large language model (LLM) agents depend on both the model and its harness, which shapes context, tool use, and interaction with the environment. Useful harness guidance changes as execution unfolds and depends on the model's ability to use it, yet existing approaches often optimize one component while fixing the other. We propose LEAH, which jointly trains an agent to decide when and how to revise its harness during execution and to act under the resulting harness toward a shared task objective. Since revisions affect task outcomes through subsequent execution, joint learning requires distinguishing their contributions. LEAH uses hierarchical reinforcement learning with separate advantage estimates: revisions receive credit for downstream execution, while actions are evaluated relative to their harness context. LEAH achieves a 96.43% success rate on ALFWorld, a task goal completion score of 65.48 on AppWorld Test-Normal, and 76.04% on MedAgentGym, outperforming execution-only and flat joint-training baselines. Further analyses show that revising the harness during execution can improve performance without updating model weights. During joint training, revisions become more relevant and actionable, and the agent makes more task progress after revising its harness.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.