Back to results

Stellenbosch : Stellenbosch University

Automatic assessment and feature-based characterisation of oral narratives for Afrikaans and isiXhosa children

Abstract

dc:description.abstract

Developing narrative and comprehension skills in early childhood is critical for later literacy and academic success. Yet, less than 20% of South African 10-year-olds can read for meaning, highlighting an urgent need for early intervention. Accurate assessment is essential for identifying children who need support, but current methods rely on human judgement, which is resource-intensive and susceptible to inconsistencies. To address this, we design an automatic scoring system for low-income preschools, where large class sizes make it challenging for teachers to identify learners requiring assistance. Our research focuses on three core goals. First, we develop an automatic scoring model prioritising predictive accuracy. Our system employs automatic speech recognition (ASR), followed by linear and logistic regression models to predict language proficiency. To reduce the feature space, we apply an L1-regularised model to the vectorised text. Linear models are chosen for their simplicity and interpretability. We experiment with different linguistic units and scaling methods, finding that word-level tokens and raw feature values yield the best performance. We then compare our final linear model with a large language model (LLM) developed in parallel work. Although the LLM achieves higher accuracy in most cases, the linear model performs competitively given its simplicity. Second, beyond predictive accuracy, we examine which features the logistic model relies on most. Assuming perfect ASR transcriptions,we conduct quantitative analyses using permutation feature importance and exploratory data analysis, focusing on proficiency features, part-of speech counts and specific keywords. We also consult speech therapists to qualitatively interpret the relevance of these features relative to established measures of narrative development. Key indicators include lexical diversity and productivity, while speech production features such as articulation rate show unexpectedly limited predictive value. Although relevant keywords and phrases vary by language, certain verbs, auxiliaries and nouns linked to core story elements consistently indicate stronger narrative skills in Afrikaans and isiXhosa. Third, in collaboration with Retief Louw, we deploy and evaluate our system through a mobile application, Auto MAIN. We collect feedback from 20 adult participants, including qualified speech-language and hearing therapists and speech therapy students. While further refinements are required, 95% of participants agree that the app has the potential to support professionals in identifying children at risk of language delays. Together, these findings demonstrate the potential of integrating ASR and predictive machine learning models to assist teachers in the early detection of developmental delays. We hope this work provides a foundation for automatic oral narrative assessment tools in low-resource languages, informing targeted interventions in foundational phase classrooms.

Degree

thesis:*
Grantor dc:publisher
Stellenbosch : Stellenbosch University
Year dc:date.issued
2026

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Sharratt, Emma Lori
Advisor dc:contributor.advisor
  • Kamper, Herman

Rights

Language dc:language.iso
en

Identifiers

dc:identifier.*
Repository record dc:identifier.uri
https://scholar.sun.ac.za/handle/10019.1/135854
OAI identifier oai:identifier
oai:scholar.sun.ac.za:10019.1/135854

Chain of custody

source
Harvested from
Stellenbosch University
Base URL
scholar.sun.ac.za/server/oai/request
Last updated
2026-07-24
Source record
OAI-PMH GetRecord
related terms
citation

Sharratt, Emma Lori. Automatic assessment and feature-based characterisation of oral narratives for Afrikaans and isiXhosa children. Stellenbosch : Stellenbosch University, 2026. https://scholar.sun.ac.za/handle/10019.1/135854