PACO: Parsimonious Compositional Control for Instruction-Based Video Detoxification
Abstract
The growth of toxic videos that violate safety policies has become a critical challenge for large media platforms. With the emergence of instruction-based video editors, a rising solution is video detoxification, which removes toxic content from a source video while preserving its utility. Such editors are typically flow-matching models whose velocity fields are trained on source videos with editing instructions. At sampling time, the instruction guides the flow trajectory by modifying the velocity field along an editing direction. However, this procedure faces three challenges. First, single-direction guidance neglects the rich concept structure inherent to the velocity space. Second, determining which directions and edit strengths to apply is nontrivial, as a violation usually involves only a few specific toxic concepts. Third, existing video editing benchmarks rarely focus on safety, so a principled evaluation suite is lacking. To address these challenges, we propose Parsimonious Compositional Control (PACO), a video detoxification framework that optimizes a sparse composition of edit directions to steer the flow trajectory of the editor. For every video to be edited, we curate a dictionary of atomic edit directions in the velocity space to model the concept structure, and steer the editor with a sparse linear combination of the atoms whose coefficients optimize a joint safety, utility, and sparsity objective. Finally, we curate PACO-2K, a video detoxification benchmark of 2,450 task instances under 14 safety categories built from 27 textual guideline sources and 31 video datasets. Across multiple video editors, PACO consistently improves safety over steering baselines while preserving source video content well.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.