acceptodds
Under review as a conference paper at ICLR 2027

Who Bears How Much Blame? Semantic Cooperative Games for Behavioral Accountability and Incentive Alignment in Agentic AI Systems

Abstract

LLM-based multi-agent systems form judgments and decisions through sequences of semantic transformations. Failures in cross-provider workflows can impose financial losses, and the allocation of responsibility can shape providers' future configuration choices. Execution logs identify the agents and messages involved, but not which local semantic behaviors affected the outcome.We propose a behavior-level accountability framework built on Semantic Cooperative Games (SCG). SCG traces backward from outcomes to recover local semantic nodes, provenance relations, generation links, and the agents that execute them. We model these links as semantic behaviors. BCAST performs counterfactual replay over behavior subsets, estimates their effects from terminal outcomes, and assigns values to behaviors and their output nodes without manual responsibility labels. These values support Behavioral Shapley and other allocation rules. Explicitly specifying the target behavior set, reference replacement, propagation rule, and outcome evaluation also places existing message-, step-, component-, and agent-level counterfactual methods in a common theoretical framework.To connect retrospective accountability with prospective incentive design, we use Behavioral Shapley allocations as penalties and study participants' choices among semantic modes, the configurations that generate behaviors. Under stable behavioral effects and a common cancellation reference, we bound the excess expected task loss of any pure-strategy equilibrium relative to the optimum. When each participant has a mode whose advantage exceeds its variation across contexts, minimizing the proposed penalties induces a unique equilibrium that achieves the lowest expected task loss among candidate configurations. Because objective ground-truth labels for behavioral responsibility are generally unavailable, we design two complementary empirical evaluations. The first examines semantic interventions, downstream replay, and the direction of behavioral effects in LLM workflows, assessing whether system replay can provide reliable automated evaluation signals for fine-grained behavioral attribution. The second tests whether penalty-minimizing mode choices improve task outcomes. The framework connects semantic explanations of completed executions, outcome responsibility allocation, and incentive evaluation for future deployments.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.