
Research Associate — Synthetic Data Generation
Edwin Onkoba
Biography
Edwin Onkoba leads the Synthetic Data Generation team at MCAAI — a research function that directly addresses one of the most fundamental constraints in African language AI development: the chronic scarcity of labelled training data for low-resource languages.
As a PhD candidate in Computer Science at Maseno University, Edwin's doctoral research explores techniques for generating high-fidelity synthetic text and speech data that can augment real-world corpora for low-resource language modelling. His work encompasses generative adversarial networks, large language model based data synthesis, and statistical augmentation approaches, evaluated rigorously against native speaker benchmarks to ensure that synthetic data maintains authentic linguistic properties.
Edwin's contributions are foundational to MCAAI's ability to scale its language datasets beyond what community data collection alone can achieve. Community data collection — while essential for authenticity and ethical grounding — is time intensive and resource constrained. Synthetic data generation allows the centre to dramatically increase the volume of training examples available for its NLP and ASR models, accelerating model development without compromising linguistic quality.
He works closely with the corpus linguistics and NLP teams to design synthetic pipelines that complement, rather than replace, real-world data collection efforts. His benchmarking methodology — which compares synthetic data quality against held-out native speaker recordings and texts — has become a standard quality assurance protocol across MCAAI's language programmes.
Research Focus
Publications
No publications on record yet.
Projects
No projects on record yet.
Collaborate with Edwin
Interested in this research? Reach out directly or explore partnership opportunities with MCAAI.
Get in Touch