acceptodds
Under review as a conference paper at ICLR 2027

FlowHarm: On the Propagation of Harmful Objectives in Agentic Workflows

Abstract

Large language model agents increasingly solve complex tasks through coordinated workflows of dependent subtasks. Evaluations of agent safety typically report task success, refusal rates, harm scores, or attack success rates for final responses, leaving the persistence and propagation of harmful objectives across intermediate workflow steps largely underexplored. We introduce FlowHarm, a benchmark and evaluation framework for studying harmful objective propagation in workflow-based coding agents. FlowHarm contains 200 software engineering tasks spanning 8 domains and 49 subcategories, with 600 controlled objective injections covering benign, explicitly harmful, and plausibly justified harmful conditions. We evaluate both static workflows and dynamically refined workflows using metrics for objective persistence, propagation breadth and depth, and conditional objective influence. Our results show that dynamic refinement is inconsistent: harmful objectives persist at an average rate of 0.56, reaching 0.88 in the highest case, while persistence and downstream propagation remain distinct properties. Even strong models exhibit substantial propagation in static workflows, with propagation score exceeding 0.4 for both GPT-5.6-Luna and Claude-Haiku-4.5. Moreover, justified harmful objectives consistently show stronger persistence and propagation than explicit ones. These findings show harmful objective propagation as an important and insufficiently studied safety risk in agentic workflows. Warning: This paper contains inappropriate, offensive and harmful content.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.