University of Illinois at Urbana-Champaign
End-to-end modeling for code-switching automatic speech recognition
Abstract
dc:descriptionThe end-to-end deep neural networks have been the start-of-the-art architecture for many tasks in the field of Automatic Speech Recognition (ASR). However, for code-switched speech, the persistent challenge of dataset scarcity is still a major problem. Given the difficulty in collecting code-switched corpus, it is noticeable that such deep neural network based systems is usually hard to reach high accuracy compared with mono-lingual ASR systems. In this study, we present an simple yet efficient end-to-end ASR system utilizing attention based encoder-decoder framework, specifically engineered to address the complexities of code-switched speech on a English and Mandarin code-switched dataset. To overcome the dataset constraints, our approach leverages attention mechanisms, enhancing the model's ability to focus on relevant linguistic features across different languages. We integrate BERT-multilingual and wav2vec 2.0 models to enrich the system's language understanding and acoustic processing capabilities. These integrations allow the model to capture nuanced language variations and phonetic subtleties inherent in code-switched speech. The results indicate a relatively low Mixed Error Rate (MER), demonstrating the model's effectiveness in decoding complex code-switched speech. Our findings shows that combining neural network architectures with sophisticated language models improves ASR systems' adaptability in multilingual settings. We also discuss the potential of incorporating syntax knowledge into language models to leverage linguistic information.
Degree
thesis:*- Name thesis:degree_name
- M.S.
- Level thesis:degree_level
- Thesis
- Discipline thesis:degree_discipline
- Electrical & Computer Engr
- Grantor
- University of Illinois at Urbana-Champaign
- Year dc:date
- 2024
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Zhang, Feiyu
- Contributors dc:contributor
-
- Hasegawa-Johnson, Mark
Subjects
dc:subject × 4Rights
dc:rights- Statement dc:rights
-
- Copyright 2024 Feiyu Zhang
- Language dc:language
- en, eng
Identifiers
dc:identifier.*- Handle dc:identifier
- https://hdl.handle.net/2142/124616