acceptodds
Under review as a conference paper at ICLR 2027

Inverse Reinforcement Learning with State-Dependent Foresight for Human Decision-Making

Abstract

Computational modeling of behavior represents individuals with a model fitted to their observed behavior and characterizes them through its parameters. Naturalistic tasks incorporate the complexity of real environments and yield high-dimensional behavioral trajectories. Inverse reinforcement learning (IRL) infers the latent reward function underlying these trajectories. Standard IRL, however, assumes effectively infinite-horizon planning, whereas human foresight is both bounded and state-dependent. Here we introduce **SC-AIRL** (**S**tate-**C**onditional Adversarial IRL), which builds these two cognitive properties into the planner. Per-depth policies share one reward but differ in how far ahead they sum it explicitly. A state-conditional router mixes them according to the current state. Across held-out episodes in pedestrian crossing and highway overtaking tasks, SC-AIRL improves imitation over Standard AIRL, raising top-1 accuracy from 0.549 to 0.652 and from 0.629 to 0.718, respectively, while more closely reproducing participant behavior in rollouts. Inferred planning depth correlates with self-reported anxiety and motor impulsivity. SC-AIRL thus reproduces human behavior more closely than Standard AIRL while providing an interpretable, model-derived measure of state-dependent foresight.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.