acceptodds
Under review as a conference paper at ICLR 2027

Verifiable Strategic Reasoning in Artificial Agents

Abstract

Strategic AI agents increasingly act in environments where decisions depend on anticipated responses of other agents, yet evaluation typically focuses on actions and outcomes rather than on whether the stated reasons for those actions can be independently examined. We introduce an Executable Strategic Rationale (ESR) representation that turns observable evidence, opponent-response predictions, counterfactual preferences, strategic intent, contingent plans, and the selected action into an external object that can be executed and audited without accessing private chains of thought. Candidate ESRs are evaluated from matched environment snapshots, allowing their claims and behavioral consequences to be checked under a common verification contract. In repeated auctions, outcome-oriented optimization improves utility most strongly, whereas verifier-guided optimization preserves the executable rationale interface and rationale-level properties do not improve uniformly with utility. Controlled perturbations of public strategic evidence systematically change regenerated rationales and downstream behavior, providing evidence of behavioral connectivity. The interface remains operational under out-of-distribution and external-auction transfer, although utility superiority does not transfer. These results separate behavioral performance from rationale auditability and provide a practical framework for externally testing strategic AI decisions.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.