Kenyan Languages Corpus for Natural Language Processing and Machine Translation
Seed grant for low-resource language technology
Overview
This UNESCO TWAS-BMBF Seed Grant project focused on building a Kenyan-language corpus for natural language processing and machine translation. It supported postgraduate researchers and advanced linguistic resource creation for core Kenyan languages, with a strong emphasis on open, useful and sustainable language data infrastructure.
Project logic
Aim
Produce annotated and curated corpora that support machine translation and NLP across Kenyan languages.
Context
A shortage of well-annotated corpora limits research and tool-building for Kenyan languages.
Tasks
Collect, annotate, and curate corpora; support postgraduate research and publish resources.
Success
Open datasets and published research that enable translation and language modelling experiments.
Focus areas
- Machine Translation
- NLP
- Corpus Linguistics
- Low-resource Languages