African Next Voices: Pilot Data Collection in Kenya
Speech and text collection for grassroots language inclusion
Overview
African Next Voices focused on pilot data collection in Kenya to create language resources for underrepresented communities. The project centered on transcription, speech recording and community engagement in Dholuo and Kalenjin contexts, supporting the development of more representative and locally grounded AI systems.
Project logic
Aim
Create pilot speech and text resources that enable downstream language technology for underrepresented Kenyan languages.
Context
Lack of representative datasets for many African languages prevents practical ML systems from being built and deployed for local communities.
Tasks
Run community recording sessions, transcribe and annotate data, and publish pilot datasets with community agreements.
Success
Pilot datasets published, local researchers trained, and demonstrator models showing improved local-language performance.
Focus areas
- Data Collection
- Speech Data
- Low-resource Languages
- Community Participation