AI Harms: Taxonomy, Dataset and Real-World Map
Abstract
Artificial intelligence (AI) harms are increasingly reported worldwide, yet no global infrastructure systematically maps their distribution. In contrast, humanitarian crises are monitored through established networks that continuously collect, structure, and visualize evidence on global maps. Evidence of AI harm, however, remains scattered across news reports. Existing AI-incident resources are also largely manually curated, focus on incidents instead of harms, and list cases without automatically locating or mapping them. As a result, chronic and infrastructural AI harms are largely missing, and the global distribution of AI harm remains invisible. To address this, we present the first automated AI Harm Map, a traceable pipeline that turns news reports into a public global map. It rests on a novel harm-centered taxonomy, anchored to 56 human-verified cases and 1,162 human-annotated harm reports. Building on this reference dataset, we scale coverage through automated retrieval, keyword-gated filtering, and LLM-based classification over a 2021–2026 news corpus, yielding 3,038 AI harm cases. To preserve the relational, temporal, and data provenance structure of these cases, we introduce PHTKG, a data provenance-aware temporal hypergraph that jointly models multi-role relationships, temporal dynamics, and evidence reliability, outperforming graph and hypergraph baselines on a chronological benchmark. On a chronological held-out benchmark of 174 events, our PHTKG outperforms the strongest of seven baselines by 4.68 AUC points. The learned representations generate the public interactive AI Harm Map, geolocating 2,845 of 3,038 cases across 81 countries; 193 cases cannot be assigned to a single country, including regional and global cases. Our map reveals that subcategory prevalence varies sharply by geography: weapons-related harms are least US-centered (35% US vs. 49%–65% in other regions) and clustered in conflict-affected regions (Iran, Palestine, Russia, and Ukraine), a pattern that holds when using only locations named in the reports. The dynamic AI Harm Map is available at https://anonymous.4open.science/w/AI_Harm_map-3991/Module_02/visualization/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.