Swahili and code-switched English-Swahili hate speech dataset in Zenodo

This study fills data scarcity gaps by curating a Swahili and code-switched English-Swahili hate speech dataset and annotating with hate class, target and language. The dataset combines Politikweli, AfriSenti, and Hate_Speech_Kenya and contains 101,014 tweets with multilingual content and hate target labels including nationality, social status, politics, disability, ethnicity, gender, religion and others.
Size

101,014 tweets with hate class, target and language labels

Format

Tweet dataset / CSV

License

CC BY

More datasets

View all
AfriVoices-KE: A Multilingual Speech Dataset for Kenyan Languages (African Next Voices)
African Next Voices is a multilingual speech dataset targeting over 3,000 hours of scripted and unscripted audio across Dholuo, Kikuyu, Somali, Kalenjin and Maasai, collected under the Kenya pilot led by the KenCorpus Consortium and funded by the Gates Foundation. Recordings span eleven domains including agriculture, healthcare, financial transactions and digital government services, collected through ethical, community-led processes with speaker-disjoint train/dev/test splits.
CommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web Data
CommonLID is a community-created language identification benchmark of web text manually annotated for language, covering 109 languages (78 with at least 100 lines of data), totalling over 350,000 annotated lines. It was built as a shared task at the Workshop on Multilingual Data Quality Signals and released under the Common Crawl terms of use for evaluation-only use.

Subscribe to our Newsletter