acceptodds
Under review as a conference paper at ICLR 2027

Operant Conditioning for Language-Model Agents

Abstract

Can language-model agents be conditioned by the consequences of their own actions? We introduce Operant Contingency Learning (OCL), which constructs policy targets from resource and burden changes while preserving their identity, direction, context, and current relevance. Across Qwen3-0.6B and Gemma-3-1B-IT with three seeds each, a controlled conditioning assessment recovers all four canonical operant effects. After 960 trials, positive and negative reinforcement increase held-out action probability by 21.50 and 17.75 percentage points relative to exposure-matched yoked controls, while positive and negative punishment decrease it by 10.16 and 9.10 points. All 24 model–seed–effect combinations change in the predicted direction relative to their calibrated initial policies, while yoked changes show no consistent direction. Separate skill-acquisition experiments use nine seeds per model on TextWorld and a simplified Telecom task derived from -bench. Under stochastic action selection, OCL-C without scalar loss achieves 77.15% error-free success on Telecom versus 59.65% for Scalar RL, and 14.55% on TextWorld versus 5.60% for Scalar RL. Both advantages occur in all 18 paired model–seed runs on each task and survive Holm correction across eight clean-success comparisons. OCL-C combined with scalar loss achieves 65.25% stochastic clean success on Telecom and 18.87% on TextWorld, revealing different observed rankings of the two OCL variants across tasks. In exploratory Telecom comparisons, OCL-C without scalar loss also reduces mean remaining burden by 45.2% relative to Scalar RL and reaches confirmed greedy mastery in 14/18 runs, compared with 11/18 for OCL-C combined with scalar loss and 5/18 for Scalar RL. During skill-acquisition, OCL-C without scalar loss exhibited mean action-probability changes in the predicted direction across all 16 task–model–consequence categories, compared with 14 for OCL-C with scalar loss and 13 for Scalar RL. These findings support typed action consequences as an effective basis for operant conditioning and task acquisition, including policy learning without a scalar task-reward loss.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.