{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/124412"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/124412","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Breaking down barriers: advancing interdisciplinary speech applications in early children’s development","abstract":"Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2024-09-16 without embargo terms","abstract_html":"Submission original under an indefinite embargo labeled &#x27;Open Access&#x27;. The submission was exported from vireo on 2024-09-16 without embargo terms","abstract_has_math":false,"creators":["Li, Jialu"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Electrical & Computer Engr","degree_department":null,"school":null,"contributors":["Hasegawa-Johnson, Mark","McElwain, Nancy L","Bhat, Suma","Varshney, Lav R"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2024,"date_issued":"2024-04-26","date_published":"2024-04-26","updated_at":"2026-07-22T22:25:00Z","subjects":["Child Vocalizations Classifications","Child-adult Speaker Diarization","Autism Diagnosis","Family Audio Analysis","Self-supervised Learning","Transfer Learning","Interdisciplinary Speech Applications"],"languages":["en","eng"],"rights":["Copyright 2024 Jialu Li"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/124412","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Hasegawa-Johnson, Mark","McElwain, Nancy L","Bhat, Suma","Varshney, Lav R"]},{"key":"dc:creator","label":"Author","values":["Li, Jialu"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2024-04-26","2024-05"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical & Computer Engr"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Child Vocalizations Classifications","Child-adult Speaker Diarization","Autism Diagnosis","Family Audio Analysis","Self-supervised Learning","Transfer Learning","Interdisciplinary Speech Applications"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2024 Jialu Li"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/124412"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2024-09-16 without embargo terms","The student, Jialu Li, accepted the attached license on 2024-04-25 at 15:58.","The student, Jialu Li, submitted this Dissertation for approval on 2024-04-25 at 16:09.","This Dissertation was approved for publication on 2024-04-26 at 10:54.","DSpace SAF Submission Ingestion Package generated from Vireo submission #20650 on 2024-09-16 at 00:37:04","This thesis aims to develop interdisciplinary speech applications using machine learning algorithms to identify children with developmental disorders or speech and language delays early. Specifically, we build machine learning models to capture critical adult-child interactions under different social contexts, including turn-taking vocalizations between parents and infants (under 14 months old) at home or joint attention between clinicians and toddlers (1-2 years old) at clinics. Turn-taking vocalizations are considered as coordinated interactions; no response and co-vocalizations are considered as uncoordinated interactions. Previous research has shown that daily repeated and reinforced uncoordinated interactions may contribute to mental health problems in children in the long run. In autism screening, detecting whether clinicians and children establish joint attention during semi-structured assessments is crucial, as this is considered a key factor for diagnosing autism. To achieve this goal, we focused on two speech-processing tasks: speaker diarization (identify who spoke when) and vocalization classifications (identify the type of vocalization given a speaker). Because annotating audio is a labor- and time-consuming task, the thesis addresses the technical difficulties in improving the performance of speech-processing models given a limited amount of labeled audio. We explore several transfer learning techniques within supervised learning as well as leverage self-supervised learning for enhancing child audio analysis tasks. With the self-supervised learning scheme, we show that the performance of proposed interdisciplinary speech applications achieved significant advancement in child audio analysis tasks. This thesis expands the application of traditional speech technology like speech-to-text and text-to-speech, exploring its potential in other disciplines such as psychology and healthcare."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Breaking down barriers: advancing interdisciplinary speech applications in early children’s development"]}]}],"canonical_facts":{"dc:contributor":["Hasegawa-Johnson, Mark","McElwain, Nancy L","Bhat, Suma","Varshney, Lav R"],"dc:creator":["Li, Jialu"],"dc:date":["2024-04-26","2024-05"],"dc:description":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2024-09-16 without embargo terms","The student, Jialu Li, accepted the attached license on 2024-04-25 at 15:58.","The student, Jialu Li, submitted this Dissertation for approval on 2024-04-25 at 16:09.","This Dissertation was approved for publication on 2024-04-26 at 10:54.","DSpace SAF Submission Ingestion Package generated from Vireo submission #20650 on 2024-09-16 at 00:37:04","This thesis aims to develop interdisciplinary speech applications using machine learning algorithms to identify children with developmental disorders or speech and language delays early. Specifically, we build machine learning models to capture critical adult-child interactions under different social contexts, including turn-taking vocalizations between parents and infants (under 14 months old) at home or joint attention between clinicians and toddlers (1-2 years old) at clinics. Turn-taking vocalizations are considered as coordinated interactions; no response and co-vocalizations are considered as uncoordinated interactions. Previous research has shown that daily repeated and reinforced uncoordinated interactions may contribute to mental health problems in children in the long run. In autism screening, detecting whether clinicians and children establish joint attention during semi-structured assessments is crucial, as this is considered a key factor for diagnosing autism. To achieve this goal, we focused on two speech-processing tasks: speaker diarization (identify who spoke when) and vocalization classifications (identify the type of vocalization given a speaker). Because annotating audio is a labor- and time-consuming task, the thesis addresses the technical difficulties in improving the performance of speech-processing models given a limited amount of labeled audio. We explore several transfer learning techniques within supervised learning as well as leverage self-supervised learning for enhancing child audio analysis tasks. With the self-supervised learning scheme, we show that the performance of proposed interdisciplinary speech applications achieved significant advancement in child audio analysis tasks. This thesis expands the application of traditional speech technology like speech-to-text and text-to-speech, exploring its potential in other disciplines such as psychology and healthcare."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/124412"],"dc:language":["en","eng"],"dc:rights":["Copyright 2024 Jialu Li"],"dc:subject":["Child Vocalizations Classifications","Child-adult Speaker Diarization","Autism Diagnosis","Family Audio Analysis","Self-supervised Learning","Transfer Learning","Interdisciplinary Speech Applications"],"dc:title":["Breaking down barriers: advancing interdisciplinary speech applications in early children’s development"],"dc:type":["text"],"thesis:degree_discipline":["Electrical & Computer Engr"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:00Z"}