acceptodds
Under review as a conference paper at ICLR 2027

Emergent foraging strategies and functional neural motifs in homeostatic Deep RL agents

Abstract

Adaptive behavior in natural environments depends on a dynamic coupling between an organism’s internal physiological state and its ongoing motor decisions. However, the computational mechanisms underlying the integration of interoceptive signals with motor control in embodied agents remain poorly understood. We train embodied artificial agents on a continuous two-resource foraging task using a homeostatic reinforcement-learning (HRL) objective, in which the agent minimizes deviation of its internal state from a set point rather than maximizing extrinsic reward. The resulting agents develop state-dependent search strategies that parallel biological foraging principles. Both exploration extent and movement velocity scale with homeostatic state, a dynamic analogous to hunger-modulated behavior. Clustering population activity around foraging events, such as food picking, reveals recurring temporal motifs: pre-event, peri-event, post-event, ramping, and sustained. Systematic cluster-level ablation demonstrates that these motifs contribute distinctly to behavior; for instance, ablating ramping clusters yields hyperactive, excessive-exploration policies. To assess the importance of learning under physiological constraints, we also train agents with an identical neural architecture, body, and environment, but using conventional reinforcement learning (RL). On the behavioral side, the RL agents lack the general state-dependent foraging strategies; and on the neural side, most of the functional neural motifs found in HRL agents become less prevalent or disappear. These results provide a neuroethologically grounded computational framework that links behavior to neural circuit dynamics and embodied constraints.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.