acceptodds
Under review as a conference paper at ICLR 2027

CoRLA: From Next-State Feedback to Long-Horizon Credit for LLM Agent Training

Abstract

Next-state feedback provides dense supervision for large language model agents, but noisy local evaluations may not reflect an action's contribution to eventual task success. We propose CoRLA (Conformal Reward and Long-horizon Attribution), an uncertainty-aware reinforcement learning framework that connects reliable reward inference with long-horizon credit assignment. CoRLA uses conformal prediction to construct reward intervals from labels supplied by a frozen LLM teacher, updating calibration online as the policy evolves. The framework learns a potential from terminal returns under these interval constraints and propagates uncertainty-gated potential differences to assign long-horizon credit. Selective on-policy distillation adds complementary token-level supervision through hindsight corrections from low-uncertainty transitions. Policy optimization integrates local progress, long-horizon credit, and token-level supervision. Experiments show that CoRLA exceeds OpenClaw-RL by 11.2 percentage points in Terminal Bench 2.0 success rate, 8.2 points in SWE-bench Verified resolved rate, and 10.0 points in AIME 24/25 accuracy. Compared with local process rewards, CoRLA's credits also align more closely with return changes under sampled action replacements.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.