KenCorpus: Kenyan Language Corpus for NLP and Machine Learning

Text and speech resources for Swahili, Dholuo and Luhya

Completed · started 2021

Overview

KenCorpus was a Lacuna Fund-supported project designed to build a Kenyan-language corpus that could support NLP and machine learning tasks across major Kenyan languages. The project created text and speech data resources for Swahili, Dholuo and Luhya and produced proof-of-concept systems for speech recognition and question answering.

Project logic

  1. Aim

    Create open text and speech resources to enable NLP and ML for Kenyan languages.

  2. Context

    Many Kenyan languages lack curated datasets needed for reliable ML models.

  3. Tasks

    Collect corpora, validate annotations, release datasets and produce demo systems.

  4. Success

    Open public-domain corpora and demonstration systems used by researchers and practitioners.

Focus areas

  • NLP
  • Speech Data
  • Machine Learning
  • Low-resource Languages

Subscribe to our Newsletter