acceptodds
Under review as a conference paper at ICLR 2027

CAPER: Closing the Loop between LLM Action Priors and Reinforcement Learning

Abstract

Reinforcement learning (RL) often requires extensive environment interaction, while large language models (LLMs) can provide action priors that guide exploration using pretrained knowledge. However, fixed LLM priors may be mismatched with the environment and cannot adapt when grounded failures reveal unsuitable guidance. We propose losed-loop ction rior nhancement for einforcement Learning (CAPER), which alternates between RL policy learning and failure-guided action-prior refinement. Low-return trajectories identify weaknesses in the current policy–prior pair, which guide the retrieval of relevant task knowledge and the generation of revised prompt candidates. Candidate prompts are selected either through direct return evaluation or through agreement with actions in high-return trajectories, trading evaluation fidelity for environment-interaction cost. Across five tasks from three environments and multiple RL settings, CAPER consistently improves over baselines under total environment-interaction budgets. The implementation is available at https://anonymous.4open.science/r/CAPER-FD27.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.