African Next Voices is a multilingual speech dataset targeting over 3,000 hours of scripted and unscripted audio across Dholuo, Kikuyu, Somali, Kalenjin and Maasai, collected under the Kenya pilot led by the KenCorpus Consortium and funded by the Gates Foundation. Recordings span eleven domains including agriculture, healthcare, financial transactions and digital government services, collected through ethical, community-led processes with speaker-disjoint train/dev/test splits.
Sign Language dataset in Zenodo
A sign language dataset hosted in Zenodo for Kenyan Sign Language research and development, supporting computer vision and assistive technology work for deaf and hard-of-hearing communities.
Size
Annotated sign language video resources
Format
Video dataset
License
CC BY
More datasets
View allCommonLID is a community-created language identification benchmark of web text manually annotated for language, covering 109 languages (78 with at least 100 lines of data), totalling over 350,000 annotated lines. It was built as a shared task at the Workshop on Multilingual Data Quality Signals and released under the Common Crawl terms of use for evaluation-only use.
PolitiKweli is the first Swahili-English code-switched dataset for political misinformation classification in Kenya. It contains 6,345 code-switched texts alongside 22,954 English texts and 211 Swahili texts, sourced from Twitter (now X) and labelled as fake, fact or neutral against a fact-checked reference set built as part of the same study.