Building NLP Text and Speech Datasets for Low Resourced Languages in East Africa

Collaborative text and speech resource development

Completed · started 2021

Overview

This project focused on building text and speech datasets for low-resource languages in East Africa, with a particular emphasis on Kiswahili and related language contexts. It was developed through collaborative work with research partners and built on open, community-conscious data practices.

Project logic

  1. Aim

    Generate reusable language resources and workflows for low-resource East African languages.

  2. Context

    Cross-border collaboration and shared resource practices are needed to scale dataset creation across the region.

  3. Tasks

    Coordinate partners, collect and annotate corpora, and publish shared datasets and guidance.

  4. Success

    Reusable datasets and published methods that inform subsequent regional datasets and research.

Focus areas

  • Speech Data
  • Text Data
  • East Africa
  • Collaborative AI

Subscribe to our Newsletter