Web(er) of Hate: A Survey on How Hate Speech Is Typed

19 June 2025

Luna Wang

Andrew Caines

Alice Hutchings

ArXiv (abs)PDF HTML Github

Main:8 Pages

2 Figures

Bibliography:8 Pages

11 Tables

Appendix:11 Pages

Abstract

The curation of hate speech datasets involves complex design decisions that balance competing priorities. This paper critically examines these methodological choices in a diverse range of datasets, highlighting common themes and practices, and their implications for dataset reliability. Drawing on Max Weber's notion of ideal types, we argue for a reflexive approach in dataset creation, urging researchers to acknowledge their own value judgments during dataset construction, fostering transparency and methodological rigour.

View on arXiv

Comments on this paper