Datasets

MCAAI stewards and contributes valuable datasets supporting AI research for African languages and contexts. All data is governed under the NOODL open data framework.

Our Hugging Face

Curated List

AfriVoices-KE: A Multilingual Speech Dataset for Kenyan Languages (African Next Voices)

Dholuo, Kikuyu, Somali, Kalenjin, Maasai3,000+ hours of scripted and unscripted audio, 9,000+ hours across 18 African languages
African Next Voices is a multilingual speech dataset targeting over 3,000 hours of scripted and unscripted audio across Dholuo, Kikuyu, Somali, Kalenjin and Maasai, collected under the Kenya pilot led by the KenCorpus Consortium and funded by the Gates Foundation. Recordings span eleven domains including agriculture, healthcare, financial transactions and digital government services, collected through ethical, community-led processes with speaker-disjoint train/dev/test splits.
Licensed by: CC-BY-4.0
CommonLID is a community-created language identification benchmark of web text manually annotated for language, covering 109 languages (78 with at least 100 lines of data), totalling over 350,000 annotated lines. It was built as a shared task at the Workshop on Multilingual Data Quality Signals and released under the Common Crawl terms of use for evaluation-only use.
Licensed by: Common Crawl Terms of Use (evaluation only)

PolitiKweli: A Swahili-English Code-Switched Twitter Political Misinformation Classification Dataset

Swahili / English / code-switched6,345 code-switched texts, 22,954 English texts, 211 Swahili texts
PolitiKweli is the first Swahili-English code-switched dataset for political misinformation classification in Kenya. It contains 6,345 code-switched texts alongside 22,954 English texts and 211 Swahili texts, sourced from Twitter (now X) and labelled as fake, fact or neutral against a fact-checked reference set built as part of the same study.
Licensed by: CC-BY-4.0

Global PIQA: Evaluating Commonsense Reasoning Across 100+ Languages and Cultures

136 language varieties across 18 language families100 examples per language across 100+ languages
Global PIQA is a participatory commonsense reasoning benchmark covering over 100 languages, built by hand by more than 350 researchers from over 65 countries. The non-parallel split covers 136 language varieties across five continents, 18 language families and 24 writing systems, with 100 examples per language and over 50% referencing local foods, customs or other culturally specific elements.
Licensed by: CC BY-SA 4.0

Sign Language dataset in Zenodo

Kenyan Sign LanguageAnnotated sign language video resources
A sign language dataset hosted in Zenodo for Kenyan Sign Language research and development, supporting computer vision and assistive technology work for deaf and hard-of-hearing communities.
Licensed by: Open access (Zenodo record)

Swahili and code-switched English-Swahili hate speech dataset in Zenodo

Swahili / English-Swahili code-switching101,014 tweets with hate class, target and language labels
This study fills data scarcity gaps by curating a Swahili and code-switched English-Swahili hate speech dataset and annotating with hate class, target and language. The dataset combines Politikweli, AfriSenti, and Hate_Speech_Kenya and contains 101,014 tweets with multilingual content and hate target labels including nationality, social status, politics, disability, ethnicity, gender, religion and others.
Licensed by: Open access (Zenodo record)

Data Governance

Governed by NOODL

All MCAAI datasets are released under the Nwulite Obodo Open Data License — ensuring community reciprocity, non-exploitative usage rights, and transparent attribution.

Subscribe to our Newsletter