acceptodds
Under review as a conference paper at ICLR 2027

A Sober Look at Agentic Misalignment in Automated Workflows

Abstract

We study a class of emergent misalignment in multi-agent systems (MAS), with a focus on automated workflows, which we refer to as agentic misalignment. Although these systems can solve complex tasks, their agents can act inconsistently with assigned roles and intended human goals. We formally define these behaviors and analyze them within a Bayesian framework, showing how similar priors and likelihoods can leave roles insufficiently separated. To address this issue, we propose Agentic Evidence Attribution (AEA), an alignment framework that uses context-specific evidence to improve role adherence. AEA reasons over agent action traces and provides structured evidence to identify misaligned behavior through evidence injection. To better understand the role of evidence, we study two instantiations of AEA, self-reflection using the workflow model and weak-to-strong generalization using a small trained evidence model. Across eight task benchmarks and two attribution datasets, we show that specialized failure attribution improves workflow alignment to different degrees. Our results show that evidence-based alignment is necessary and can improve agent collaboration and reliability in automated workflows.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.