acceptodds
Under review as a conference paper at ICLR 2027

CogRL: Enhancing Cognitive Reasoning of LLM Agents via Reinforcement Learning

Abstract

Recent advances in large language models (LLMs) have enhanced agents’ reasoning abilities, but deploying them in dynamic environments remains challenging due to misalignment between internal reasoning and environmental dynamics. Existing RL-based methods often couple reasoning with action and rely only on sparse task rewards, leading to ungrounded reasoning and poor sample efficiency—especially in multi-round interactions. We propose CogRL, an integrated RL framework for enhancing two-stage cognitive reasoning. CogRL makes reasoning outcomes verifiable through a Subgoal Reasoning Policy (SRP) that predicts intended next states, supported by an Action Reasoning Policy (ARP) that generates actions to realize these transitions. Its integrated policy optimization combines separately normalized task-return and state-consistency advantages to train the SRP, enabling learning from unsuccessful trajectories even when task rewards are uninformative. The ARP is trained using periodic inverse-dynamics to support stable online learning. Experiments across multiple interactive environments demonstrate improved task success and data efficiency over strong baselines, with further analysis showing reduced state hallucinations. Our code is available at \url{https://anonymous.4open.science/r/CogRL

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.