CoRLA: From Next-State Feedback to Long-Horizon Credit for LLM Agent Training
Abstract
Next-state feedback provides dense supervision for large language model agents, but noisy local evaluations may not reflect an action's contribution to eventual task success. We propose CoRLA (Conformal Reward and Long-horizon Attribution), an uncertainty-aware reinforcement learning framework that connects reliable reward inference with long-horizon credit assignment. CoRLA uses conformal prediction to construct reward intervals from labels supplied by a frozen LLM teacher, updating calibration online as the policy evolves. The framework learns a potential from terminal returns under these interval constraints and propagates uncertainty-gated potential differences to assign long-horizon credit. Selective on-policy distillation adds complementary token-level supervision through hindsight corrections from low-uncertainty transitions. Policy optimization integrates local progress, long-horizon credit, and token-level supervision. Experiments show that CoRLA exceeds OpenClaw-RL by 11.2 percentage points in Terminal Bench 2.0 success rate, 8.2 points in SWE-bench Verified resolved rate, and 10.0 points in AIME 24/25 accuracy. Compared with local process rewards, CoRLA's credits also align more closely with return changes under sampled action replacements.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.