acceptodds
Under review as a conference paper at ICLR 2027

TRACE: Capability-Targeted Agentic Training

Abstract

Models fail to complete agentic tasks often due to missing core capabilities. However, mainstream approaches for addressing these failures often rely on fine-tuning directly on the target environments or generating synthetic data that is not targeted to the model’s actual capability deficits, resulting in low sample efficiency and limited generalizability. We introduce TRACE (Turning Recurrent Agent failures into Capability targeted training Environments), an end-to-end system for environment specific agent self-improvement. TRACE contrasts successful and failed trajectories to automatically identify lacking capabilities, synthesizes a targeted training environment for each that rewards whether the capability was exercised, trains a LoRA adapter via RL on each synthetic environment, and then trains a mixture-of-experts (MoE) model over the capability adapters. TRACE can be effectively applied across different environments, improving over the base agent by +15.3 points on -Bench (customer service) and +15 points Pass@1 on SWE-bench Verified (software engineering), outperforming the strongest external baselines, GEPA and SWE-RL, by +8.6 points and +8.4 points, respectively. In addition, TRACE is more sample-efficient than strong finetuning baselines: using one-fourth the number of rollouts, TRACE outperforms the best-performing baselines, GRPO and GEPA, and achieves higher final accuracy by +10.4 and +8.6 points on -Bench.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.