Stellenbosch : Stellenbosch University
Building Speech-Driven Assessment Tools for Afrikaans and isiXhosa Children
Abstract
dc:description.abstractIlliteracy is a root cause of many socio-economic problems in South Africa. However, problems with reading often start before a child enters school. It is therefore essential to identify children with development problems early to allow for effective interventions. However, current language assessment tools are not designed for local languages, require trained administrators and lack the scalability needed for widespread screening in resource-constrained settings. We address this challenge by developing automated, speech-driven assessment tools for Afrikaans and isiXhosa. This thesis makes four contributions. First, we develop Mamela, an automatic speech recognition (ASR) system that we optimise for child speech for these lowresource languages. Second, we present an automated pipeline for language sample analysis (LSA) from oral child-adult interactions, which calculates key clinical metrics of language development. Third, we examine the multilingual assessment instrument for narratives (MAIN), a tool used by speech therapists to assess children’s development through oral narratives. We integrate Mamela into a pipeline with speaker diarisation, ASR, machine translation and a large language model (LLM) to predict the MAIN scores automatically. Fourth, we integrate the developed MAIN pipeline into a user-friendly mobile application for speech therapists and educators. For each contribution, we perform a corresponding evaluation. First, our ASR fine-tuning strategy improves the word error rates for Afrikaans by 72% and for isiXhosa by 61% relative to the baseline models. The resulting error rates, however, remain high compared to other high-resource languages. Second, the outputs of our automated LSA pipeline achieve an R2 value of 0.91 for the mean length of utterance and 0.95 for the total number of words. Third, our LLM-based classifier identifies children requiring intervention with an accuracy of over 80% for Afrikaans and 64% for isiXhosa. This performance is on par with human assessors despite using the imperfect ASR transcripts, suggesting that such end-to-end systems are viable in challenging acoustic conditions. Fourth, an expert panel of speech therapists and students evaluated the MAIN mobile application, showing that the application can potentially reduce the assessment time from 30 to 45 minutes to under one minute. Ninety-five percent of the panel agreed that the application can assist professionals in identifying children at risk of language development delay. This thesis demonstrates that integrating modern ASR and LLM technology can automate complex language assessments for low-resource languages. Our work translates this research into a practical tool, offering educators and speech therapists a scalable, accurate and low-effort method to identify at-risk children.
Degree
thesis:*- Grantor dc:publisher
- Stellenbosch : Stellenbosch University
- Year dc:date.issued
- 2026
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Louw, Retief Gideon Arno
- Advisors dc:contributor.advisor
-
- Kamper, H.
- Strauss, Johannes
Rights
- Language dc:language.iso
- en
Identifiers
dc:identifier.*- Repository record dc:identifier.uri
- https://scholar.sun.ac.za/handle/10019.1/136280
- OAI identifier oai:identifier
- oai:scholar.sun.ac.za:10019.1/136280