Reconstruct, Then Evolve: Grounded Data Synthesis for Harmful Meme Detection
Abstract
Harmful memes often convey abuse through implicit interactions between images and text, making reliable detection critically dependent on training data that both faithfully captures harmful semantics and covers diverse ways those semantics can be expressed. Existing work has largely improved model architectures, contextual reasoning, or inference procedures while leaving the training distribution unchanged. This leaves two data bottlenecks: Weak cross-modal grounding within individual examples and limited coverage of harmful realizations in finite training sets. Effective synthesis therefore requires preserving *what makes a meme harmful* while systematically varying *how that meaning is realized*. We introduce *Reconstruct-then-Evolve*, a data-centric framework that first reconstructs source memes into semantically grounded realizations and then expands them through verifier-guided Monte Carlo Tree Search over alternative expression strategies, scenes, templates, and visual styles. Without modifying the detector architecture or inference procedure, our synthesized data raise accuracy by 5.1-10.7 percentage points across four benchmarks and outperform prior published baselines on all four. The gains persist across multiple detector backbones, with improvements of 7.4-10.7 percentage points over fine-tuning on the original training data. These results show that systematically improving the training distribution provides a powerful and complementary path toward more robust harmful meme detection.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.