KenCorpus: Kenyan Language Corpus for NLP and Machine Learning
Text and speech resources for Swahili, Dholuo and Luhya
Overview
KenCorpus was a Lacuna Fund-supported project designed to build a Kenyan-language corpus that could support NLP and machine learning tasks across major Kenyan languages. The project created text and speech data resources for Swahili, Dholuo and Luhya and produced proof-of-concept systems for speech recognition and question answering.
Project logic
Aim
Create open text and speech resources to enable NLP and ML for Kenyan languages.
Context
Many Kenyan languages lack curated datasets needed for reliable ML models.
Tasks
Collect corpora, validate annotations, release datasets and produce demo systems.
Success
Open public-domain corpora and demonstration systems used by researchers and practitioners.
Focus areas
- NLP
- Speech Data
- Machine Learning
- Low-resource Languages