SAME SCORE, DIFFERENT OUTCOMES: REWARD INTERFACES SHAPE IN-CONTEXT REINFORCEMENT LEARNING
Abstract
Large language models can adapt across interaction rounds from feedback retained in context without updating their parameters. However, existing studies typically treat reward as a scalar signal, leaving open whether the language interface through which that signal enters the context affects multi-round adaptation. To study this question, we develop a two-stage evaluation framework based on reward-interface presentation packages, which map an underlying scalar reward into contextualized language feedback. Stage 1 examines semantic enrichment, asking whether adding an explicit feedback contract and qualitative state descriptors to a scalar reward changes adaptation. Stage 2 examines rhetorical presentation framing, holding the underlying reward and semantic scaffold structure fixed while expressing the feedback through different semantic frames: care-, competition-, survival-, trust-, and navigation-oriented frames. We conduct systematic multi-round comparisons across 16,800 twenty-round episodes on two deterministic tasks, FEVER fact verification and SQuAD 2.0 extractive question answering, using GPT-5.4-mini, Qwen3-8B, and Gemma-4-E4B-it. The results show that the effect of semantic enrichment depends on the task and model: it increases late FEVER reward by 0.474 points for GPT-5.4-mini and 0.909 points for Qwen3-8B, but provides limited or mixed benefit on SQuAD 2.0. Even with the underlying reward and scaffold structure held fixed, different rhetorical frames produce distinct multi-round adaptation trajectories and endpoint differences of up to 1.251 points. These findings suggest that the linguistic presentation of reward participates in shaping model behavior during ICRL, rather than serving merely as a textual format for transmitting a scalar signal.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.