RGFA: Reliability Guided Feedback Adaptation in Large Language Models via Selective Feedback Utilization
Abstract
Learning from feedback has long been studied in psychology as a fundamental mechanism underlying adaptive decision-making. In humans, this mechanism enables efficient adaptation even when feedback is noisy or uncertain. As large language models (LLMs) are increasingly deployed in interactive and decision-making tasks, a fundamental question arises: Can LLMs learn from feedback as effectively as humans, and how do their feedback-driven learning behaviors differ? Using a controlled probabilistic reversal learning paradigm, we compare how LLMs and humans learn from feedback under uncertainty. We find that LLMs perform comparably to humans with positive feedback but are much less effective with negative feedback. This gap arises mainly because LLMs struggle to distinguish informative negative feedback from environmental noise and to use it to improve their policies. To address these limitations, we propose \underline Reliability Guided Feedback Adaptation (RGFA), a lightweight approach that regulates policy updates according to the reliability of accumulated feedback evidence. Experiments show that RGFA substantially improves learning from negative feedback, and thus narrows the behavioral gap between LLMs and humans.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.