{"id":{"repo_id":"stellenbosch","oai_identifier":"oai:scholar.sun.ac.za:10019.1/135854"},"canonical_url":"https://search.dev.ndltd.org/etd/stellenbosch/oai:scholar.sun.ac.za:10019.1/135854","repository":{"repo_id":"stellenbosch","name":"Stellenbosch University","base_url":"https://scholar.sun.ac.za/server/oai/request"},"display":{"title":"Automatic assessment and feature-based characterisation of oral narratives for Afrikaans and isiXhosa children","abstract":"Developing narrative and comprehension skills in early childhood is critical for later literacy and academic success. Yet, less than 20% of South African 10-year-olds can read for meaning, highlighting an urgent need for early intervention. Accurate assessment is essential for identifying children who need support, but current methods rely on human judgement, which is resource-intensive and susceptible to inconsistencies. To address this, we design an automatic scoring system for low-income preschools, where large class sizes make it challenging for teachers to identify learners requiring assistance. Our research focuses on three core goals. First, we develop an automatic scoring model prioritising predictive accuracy. Our system employs automatic speech recognition (ASR), followed by linear and logistic regression models to predict language proficiency. To reduce the feature space, we apply an L1-regularised model to the vectorised text. Linear models are chosen for their simplicity and interpretability. We experiment with different linguistic units and scaling methods, finding that word-level tokens and raw feature values yield the best performance. We then compare our final linear model with a large language model (LLM) developed in parallel work. Although the LLM achieves higher accuracy in most cases, the linear model performs competitively given its simplicity. Second, beyond predictive accuracy, we examine which features the logistic model relies on most. Assuming perfect ASR transcriptions,we conduct quantitative analyses using permutation feature importance and exploratory data analysis, focusing on proficiency features, part-of speech counts and specific keywords. We also consult speech therapists to qualitatively interpret the relevance of these features relative to established measures of narrative development. Key indicators include lexical diversity and productivity, while speech production features such as articulation rate show unexpectedly limited predictive value. Although relevant keywords and phrases vary by language, certain verbs, auxiliaries and nouns linked to core story elements consistently indicate stronger narrative skills in Afrikaans and isiXhosa. Third, in collaboration with Retief Louw, we deploy and evaluate our system through a mobile application, Auto MAIN. We collect feedback from 20 adult participants, including qualified speech-language and hearing therapists and speech therapy students. While further refinements are required, 95% of participants agree that the app has the potential to support professionals in identifying children at risk of language delays. Together, these findings demonstrate the potential of integrating ASR and predictive machine learning models to assist teachers in the early detection of developmental delays. We hope this work provides a foundation for automatic oral narrative assessment tools in low-resource languages, informing targeted interventions in foundational phase classrooms.","abstract_html":"Developing narrative and comprehension skills in early childhood is critical for later literacy and academic success. Yet, less than 20% of South African 10-year-olds can read for meaning, highlighting an urgent need for early intervention. Accurate assessment is essential for identifying children who need support, but current methods rely on human judgement, which is resource-intensive and susceptible to inconsistencies. To address this, we design an automatic scoring system for low-income preschools, where large class sizes make it challenging for teachers to identify learners requiring assistance. Our research focuses on three core goals. First, we develop an automatic scoring model prioritising predictive accuracy. Our system employs automatic speech recognition (ASR), followed by linear and logistic regression models to predict language proficiency. To reduce the feature space, we apply an L1-regularised model to the vectorised text. Linear models are chosen for their simplicity and interpretability. We experiment with different linguistic units and scaling methods, finding that word-level tokens and raw feature values yield the best performance. We then compare our final linear model with a large language model (LLM) developed in parallel work. Although the LLM achieves higher accuracy in most cases, the linear model performs competitively given its simplicity. Second, beyond predictive accuracy, we examine which features the logistic model relies on most. Assuming perfect ASR transcriptions,we conduct quantitative analyses using permutation feature importance and exploratory data analysis, focusing on proficiency features, part-of speech counts and specific keywords. We also consult speech therapists to qualitatively interpret the relevance of these features relative to established measures of narrative development. Key indicators include lexical diversity and productivity, while speech production features such as articulation rate show unexpectedly limited predictive value. Although relevant keywords and phrases vary by language, certain verbs, auxiliaries and nouns linked to core story elements consistently indicate stronger narrative skills in Afrikaans and isiXhosa. Third, in collaboration with Retief Louw, we deploy and evaluate our system through a mobile application, Auto MAIN. We collect feedback from 20 adult participants, including qualified speech-language and hearing therapists and speech therapy students. While further refinements are required, 95% of participants agree that the app has the potential to support professionals in identifying children at risk of language delays. Together, these findings demonstrate the potential of integrating ASR and predictive machine learning models to assist teachers in the early detection of developmental delays. We hope this work provides a foundation for automatic oral narrative assessment tools in low-resource languages, informing targeted interventions in foundational phase classrooms.","abstract_has_math":false,"creators":["Sharratt, Emma Lori"],"institution":"Stellenbosch : Stellenbosch University","degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":["Kamper, Herman"],"committee_chairs":[],"committee_members":[],"year":2026,"date_issued":"2026-03","date_published":"2026-03","updated_at":"2026-07-24T04:40:06Z","subjects":[],"languages":["en"],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://scholar.sun.ac.za/handle/10019.1/135854","outbound_label":"Repository record","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Kamper, Herman"]},{"key":"dc:contributor.other","label":"Dc Contributor Other","values":["Stellenbosch University. Faculty of Engineering. Dept. of Electrical and Electronic Engineering."]},{"key":"dc:creator","label":"Author","values":["Sharratt, Emma Lori"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2026-04-13T12:27:26Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2026-04-13T12:27:26Z"]},{"key":"dc:date.issued","label":"Date","values":["2026-03"]},{"key":"dc:publisher","label":"Institution","values":["Stellenbosch : Stellenbosch University"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["en"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://scholar.sun.ac.za/handle/10019.1/135854"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Thesis (MEng)--Stellenbosch University, 2026.","Sharratt, E. L. 2026. Automatic assessment and feature-based characterisation of oral narratives for Afrikaans and isiXhosa children. Unpublished masters thesis. Stellenbosch: Stellenbosch University [online]. Available: https://scholar.sun.ac.za/items/1d249289-3e1a-4d3f-a928-a181faaaf004"]},{"key":"dc:description.abstract","label":"Abstract","values":["Developing narrative and comprehension skills in early childhood is critical for later literacy and academic success. Yet, less than 20% of South African 10-year-olds can read for meaning, highlighting an urgent need for early intervention. Accurate assessment is essential for identifying children who need support, but current methods rely on human judgement, which is resource-intensive and susceptible to inconsistencies. To address this, we design an automatic scoring system for low-income preschools, where large class sizes make it challenging for teachers to identify learners requiring assistance. Our research focuses on three core goals. First, we develop an automatic scoring model prioritising predictive accuracy. Our system employs automatic speech recognition (ASR), followed by linear and logistic regression models to predict language proficiency. To reduce the feature space, we apply an L1-regularised model to the vectorised text. Linear models are chosen for their simplicity and interpretability. We experiment with different linguistic units and scaling methods, finding that word-level tokens and raw feature values yield the best performance. We then compare our final linear model with a large language model (LLM) developed in parallel work. Although the LLM achieves higher accuracy in most cases, the linear model performs competitively given its simplicity. Second, beyond predictive accuracy, we examine which features the logistic model relies on most. Assuming perfect ASR transcriptions,we conduct quantitative analyses using permutation feature importance and exploratory data analysis, focusing on proficiency features, part-of speech counts and specific keywords. We also consult speech therapists to qualitatively interpret the relevance of these features relative to established measures of narrative development. Key indicators include lexical diversity and productivity, while speech production features such as articulation rate show unexpectedly limited predictive value. Although relevant keywords and phrases vary by language, certain verbs, auxiliaries and nouns linked to core story elements consistently indicate stronger narrative skills in Afrikaans and isiXhosa. Third, in collaboration with Retief Louw, we deploy and evaluate our system through a mobile application, Auto MAIN. We collect feedback from 20 adult participants, including qualified speech-language and hearing therapists and speech therapy students. While further refinements are required, 95% of participants agree that the app has the potential to support professionals in identifying children at risk of language delays. Together, these findings demonstrate the potential of integrating ASR and predictive machine learning models to assist teachers in the early detection of developmental delays. We hope this work provides a foundation for automatic oral narrative assessment tools in low-resource languages, informing targeted interventions in foundational phase classrooms."]},{"key":"dc:title","label":"Title","values":["Automatic assessment and feature-based characterisation of oral narratives for Afrikaans and isiXhosa children"]}]}],"canonical_facts":{"dc:contributor.advisor":["Kamper, Herman"],"dc:contributor.other":["Stellenbosch University. Faculty of Engineering. Dept. of Electrical and Electronic Engineering."],"dc:creator":["Sharratt, Emma Lori"],"dc:date.accessioned":["2026-04-13T12:27:26Z"],"dc:date.available":["2026-04-13T12:27:26Z"],"dc:date.issued":["2026-03"],"dc:description":["Thesis (MEng)--Stellenbosch University, 2026.","Sharratt, E. L. 2026. Automatic assessment and feature-based characterisation of oral narratives for Afrikaans and isiXhosa children. Unpublished masters thesis. Stellenbosch: Stellenbosch University [online]. Available: https://scholar.sun.ac.za/items/1d249289-3e1a-4d3f-a928-a181faaaf004"],"dc:description.abstract":["Developing narrative and comprehension skills in early childhood is critical for later literacy and academic success. Yet, less than 20% of South African 10-year-olds can read for meaning, highlighting an urgent need for early intervention. Accurate assessment is essential for identifying children who need support, but current methods rely on human judgement, which is resource-intensive and susceptible to inconsistencies. To address this, we design an automatic scoring system for low-income preschools, where large class sizes make it challenging for teachers to identify learners requiring assistance. Our research focuses on three core goals. First, we develop an automatic scoring model prioritising predictive accuracy. Our system employs automatic speech recognition (ASR), followed by linear and logistic regression models to predict language proficiency. To reduce the feature space, we apply an L1-regularised model to the vectorised text. Linear models are chosen for their simplicity and interpretability. We experiment with different linguistic units and scaling methods, finding that word-level tokens and raw feature values yield the best performance. We then compare our final linear model with a large language model (LLM) developed in parallel work. Although the LLM achieves higher accuracy in most cases, the linear model performs competitively given its simplicity. Second, beyond predictive accuracy, we examine which features the logistic model relies on most. Assuming perfect ASR transcriptions,we conduct quantitative analyses using permutation feature importance and exploratory data analysis, focusing on proficiency features, part-of speech counts and specific keywords. We also consult speech therapists to qualitatively interpret the relevance of these features relative to established measures of narrative development. Key indicators include lexical diversity and productivity, while speech production features such as articulation rate show unexpectedly limited predictive value. Although relevant keywords and phrases vary by language, certain verbs, auxiliaries and nouns linked to core story elements consistently indicate stronger narrative skills in Afrikaans and isiXhosa. Third, in collaboration with Retief Louw, we deploy and evaluate our system through a mobile application, Auto MAIN. We collect feedback from 20 adult participants, including qualified speech-language and hearing therapists and speech therapy students. While further refinements are required, 95% of participants agree that the app has the potential to support professionals in identifying children at risk of language delays. Together, these findings demonstrate the potential of integrating ASR and predictive machine learning models to assist teachers in the early detection of developmental delays. We hope this work provides a foundation for automatic oral narrative assessment tools in low-resource languages, informing targeted interventions in foundational phase classrooms."],"dc:identifier.uri":["https://scholar.sun.ac.za/handle/10019.1/135854"],"dc:language.iso":["en"],"dc:publisher":["Stellenbosch : Stellenbosch University"],"dc:title":["Automatic assessment and feature-based characterisation of oral narratives for Afrikaans and isiXhosa children"],"dc:type":["Thesis"]},"updated_at":"2026-07-24T04:40:06Z"}