BlameShift: Blame Integrity of Multi-Agent Failure Attribution under Ownership-Constrained Narrative Edits
Abstract
Multi-agent failure attribution relies on mixed logs combining trusted execution events with agent-authored messages. An agent responsible for a failure can exploit this interface by rewriting its own narrative to shift blame onto an innocent peer while execution records remain fixed. We define blame integrity as the invariance of an audit interface–judge mechanism to ownership-constrained narrative edits, precluding profitable deviations for blame-avoiding agents. On multi-agent benchmarks, ownership-legal accusations reliably steer flat LLM judges toward preselected victims over matched controls, an effect driven by peer accusations rather than text length, name priming, or factual fabrication. Vulnerability varies across judges, belonging to the audit mechanism rather than the text alone. Theoretically, evidence-independent style filters are either bypassable or own-blame inert, and empirically, filter-aware rewrites circumvent them. While narrative deletion trades off clean accuracy and citation verification limits steering, offline committees fail to Pareto-dominate their strongest constituent.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.