Building NLP Text and Speech Datasets for Low Resourced Languages in East Africa
Collaborative text and speech resource development
Overview
This project focused on building text and speech datasets for low-resource languages in East Africa, with a particular emphasis on Kiswahili and related language contexts. It was developed through collaborative work with research partners and built on open, community-conscious data practices.
Project logic
Aim
Generate reusable language resources and workflows for low-resource East African languages.
Context
Cross-border collaboration and shared resource practices are needed to scale dataset creation across the region.
Tasks
Coordinate partners, collect and annotate corpora, and publish shared datasets and guidance.
Success
Reusable datasets and published methods that inform subsequent regional datasets and research.
Focus areas
- Speech Data
- Text Data
- East Africa
- Collaborative AI