{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/124616"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/124616","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"End-to-end modeling for code-switching automatic speech recognition","abstract":"Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2026-05-01","abstract_html":"Submission published under a 24 month embargo labeled &#x27;Closed Access&#x27;, the embargo will last until 2026-05-01","abstract_has_math":false,"creators":["Zhang, Feiyu"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Electrical & Computer Engr","degree_department":null,"school":null,"contributors":["Hasegawa-Johnson, Mark"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2024,"date_issued":"2024-05","date_published":"2024-05","updated_at":"2026-07-22T22:25:02Z","subjects":["Speech Recognition","Code-switching","End-to-end","Embeddings"],"languages":["en","eng"],"rights":["Copyright 2024 Feiyu Zhang"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/124616","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Hasegawa-Johnson, Mark"]},{"key":"dc:creator","label":"Author","values":["Zhang, Feiyu"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2024-05","2024-05-01"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical & Computer Engr"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Speech Recognition","Code-switching","End-to-end","Embeddings"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2024 Feiyu Zhang"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/124616"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2026-05-01","The student, Feiyu Zhang, accepted the attached license on 2024-04-30 at 18:35.","The student, Feiyu Zhang, submitted this Thesis for approval on 2024-04-30 at 18:44.","This Thesis was approved for publication on 2024-05-01 at 09:17.","DSpace SAF Submission Ingestion Package generated from Vireo submission #20176 on 2024-09-16 at 00:48:13","The end-to-end deep neural networks have been the start-of-the-art architecture for many tasks in the field of Automatic Speech Recognition (ASR). However, for code-switched speech, the persistent challenge of dataset scarcity is still a major problem. Given the difficulty in collecting code-switched corpus, it is noticeable that such deep neural network based systems is usually hard to reach high accuracy compared with mono-lingual ASR systems. In this study, we present an simple yet efficient end-to-end ASR system utilizing attention based encoder-decoder framework, specifically engineered to address the complexities of code-switched speech on a English and Mandarin code-switched dataset. To overcome the dataset constraints, our approach leverages attention mechanisms, enhancing the model's ability to focus on relevant linguistic features across different languages. We integrate BERT-multilingual and wav2vec 2.0 models to enrich the system's language understanding and acoustic processing capabilities. These integrations allow the model to capture nuanced language variations and phonetic subtleties inherent in code-switched speech. The results indicate a relatively low Mixed Error Rate (MER), demonstrating the model's effectiveness in decoding complex code-switched speech. Our findings shows that combining neural network architectures with sophisticated language models improves ASR systems' adaptability in multilingual settings. We also discuss the potential of incorporating syntax knowledge into language models to leverage linguistic information."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["End-to-end modeling for code-switching automatic speech recognition"]}]}],"canonical_facts":{"dc:contributor":["Hasegawa-Johnson, Mark"],"dc:creator":["Zhang, Feiyu"],"dc:date":["2024-05","2024-05-01"],"dc:description":["Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2026-05-01","The student, Feiyu Zhang, accepted the attached license on 2024-04-30 at 18:35.","The student, Feiyu Zhang, submitted this Thesis for approval on 2024-04-30 at 18:44.","This Thesis was approved for publication on 2024-05-01 at 09:17.","DSpace SAF Submission Ingestion Package generated from Vireo submission #20176 on 2024-09-16 at 00:48:13","The end-to-end deep neural networks have been the start-of-the-art architecture for many tasks in the field of Automatic Speech Recognition (ASR). However, for code-switched speech, the persistent challenge of dataset scarcity is still a major problem. Given the difficulty in collecting code-switched corpus, it is noticeable that such deep neural network based systems is usually hard to reach high accuracy compared with mono-lingual ASR systems. In this study, we present an simple yet efficient end-to-end ASR system utilizing attention based encoder-decoder framework, specifically engineered to address the complexities of code-switched speech on a English and Mandarin code-switched dataset. To overcome the dataset constraints, our approach leverages attention mechanisms, enhancing the model's ability to focus on relevant linguistic features across different languages. We integrate BERT-multilingual and wav2vec 2.0 models to enrich the system's language understanding and acoustic processing capabilities. These integrations allow the model to capture nuanced language variations and phonetic subtleties inherent in code-switched speech. The results indicate a relatively low Mixed Error Rate (MER), demonstrating the model's effectiveness in decoding complex code-switched speech. Our findings shows that combining neural network architectures with sophisticated language models improves ASR systems' adaptability in multilingual settings. We also discuss the potential of incorporating syntax knowledge into language models to leverage linguistic information."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/124616"],"dc:language":["en","eng"],"dc:rights":["Copyright 2024 Feiyu Zhang"],"dc:subject":["Speech Recognition","Code-switching","End-to-end","Embeddings"],"dc:title":["End-to-end modeling for code-switching automatic speech recognition"],"dc:type":["text"],"thesis:degree_discipline":["Electrical & Computer Engr"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:02Z"}