acceptodds
Under review as a conference paper at ICLR 2027

Safe Preference Identification through Strategic Interaction

Abstract

Learning a partner's preferences through strategic interaction can fail for two distinct reasons: safety constraints may remove the interactions that distinguish preferences, while a misspecified behavioral model may interpret informative responses in the wrong direction. We formalize these bottlenecks through an available-information rate and a learned-evidence rate . Their separation exposes a central failure regime, but , in which the admissible environment contains identifying information but the learner cannot convert it into correctly directed evidence. We characterize constraint-induced non-identifiability, derive information-theoretic sample-complexity lower bounds governed by , and show that signed learned evidence determines asymptotic posterior drift under likelihood misspecification. Experiments in an analytic instance, an exactly enumerable strategic game, and continuous HighwayEnv merging show that future commitments can change current responses, trajectory-level evidence can improve identification over one-step evidence, and natural likelihood mismatch can reverse learned evidence. Safe strategic preference identification therefore requires both informative admissible interactions and a behavioral model that interprets their responses faithfully.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.