African Next Voices: Pilot Data Collection in Kenya

Speech and text collection for grassroots language inclusion

Completed · started 2024

Overview

African Next Voices focused on pilot data collection in Kenya to create language resources for underrepresented communities. The project centered on transcription, speech recording and community engagement in Dholuo and Kalenjin contexts, supporting the development of more representative and locally grounded AI systems.

Project logic

  1. Aim

    Create pilot speech and text resources that enable downstream language technology for underrepresented Kenyan languages.

  2. Context

    Lack of representative datasets for many African languages prevents practical ML systems from being built and deployed for local communities.

  3. Tasks

    Run community recording sessions, transcribe and annotate data, and publish pilot datasets with community agreements.

  4. Success

    Pilot datasets published, local researchers trained, and demonstrator models showing improved local-language performance.

Focus areas

  • Data Collection
  • Speech Data
  • Low-resource Languages
  • Community Participation

Subscribe to our Newsletter