Kenyan Languages Corpus for Natural Language Processing and Machine Translation

Seed grant for low-resource language technology

Completed · started 2023

Overview

This UNESCO TWAS-BMBF Seed Grant project focused on building a Kenyan-language corpus for natural language processing and machine translation. It supported postgraduate researchers and advanced linguistic resource creation for core Kenyan languages, with a strong emphasis on open, useful and sustainable language data infrastructure.

Project logic

  1. Aim

    Produce annotated and curated corpora that support machine translation and NLP across Kenyan languages.

  2. Context

    A shortage of well-annotated corpora limits research and tool-building for Kenyan languages.

  3. Tasks

    Collect, annotate, and curate corpora; support postgraduate research and publish resources.

  4. Success

    Open datasets and published research that enable translation and language modelling experiments.

Focus areas

  • Machine Translation
  • NLP
  • Corpus Linguistics
  • Low-resource Languages

Subscribe to our Newsletter